GPUs, particularly NVIDIA's H100 and B200, excel in matrix multiplication with specialized cores and high memory bandwidth, making them versatile for large language models (LLMs) despite their origins in graphics processing. Each H100 GPU boasts 990 bf16 TFLOPs/s and 3.35 TB/s memory bandwidth, while the B200 offers 2250 bf16 TFLOPs/s and 9 TB/s bandwidth, enhancing performance for ML tasks.
Weaponizing image scaling against production AI systems
Image scaling attacks exploit vulnerabilities in AI systems by using downscaled images to reveal hidden prompt injections, enabling data exfiltration from platforms like Google Gemini CLI and Vertex AI Studio.
SK hynix dethrones Samsung as world’s top DRAM maker
SK hynix has overtaken Samsung as the world's leading DRAM manufacturer, marking a significant shift in the semiconductor industry after more than 30 years of Samsung's dominance.
"AI first" and the Bus Factor of 0
The "Bus Factor" measures the risk of losing critical project knowledge, and with the rise of AI-first approaches, many teams now operate with a Bus Factor of zero, relying solely on AI-generated code without human understanding.
Show HN: Luminal – Open-source, search-based GPU compiler
Luminal is a deep learning library that leverages search-based compilation to enhance performance, aiming to be the fastest ML framework across devices, with a focus on simplicity and minimalism in its architecture.
Apple Watch wearable foundation model
Foundation models utilizing behavioral data from wearables can significantly enhance health predictions, leveraging over 2.5B hours of data from 162K individuals to optimize model architectures and tokenization strategies.
Learning about GPUs through measuring memory bandwidth
Measuring memory bandwidth in GPUs reveals significant performance differences across architectures, with the Qualcomm Adreno 740 achieving a bandwidth of over 130 GiB/s when using textures instead of buffers, highlighting the importance of resource selection in software optimization.
DeepSeek-v3.1 Release
DeepSeek-V3.1 introduces hybrid inference with two modes—Think and Non-Think—enhancing both speed and capability in agent tasks.
In a first, Google has released data on how much energy an AI prompt uses
Google’s Gemini AI prompts consume a median of 0.24 watt-hours of electricity, equivalent to running a microwave for about one second, marking a significant step in transparency regarding AI energy usage.
NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
Nemotron-Nano-9B-v2 is a hybrid Mamba-Transformer model that enhances throughput for reasoning tasks while maintaining state-of-the-art accuracy against similarly-sized models, achieving up to 6x higher inference throughput in specific settings.
Into the Omniverse: How OpenUSD and Digital Twins Are Powering Industrial and Physical AI
OpenUSD and digital twins are revolutionizing industrial and physical AI by enabling the rapid creation of physically accurate simulations, which enhance the training of AI agents and autonomous systems.
Gearing Up for the Gigawatt Data Center Age
AI factories are transforming data centers into high-performance computing units, utilizing tens to hundreds of thousands of GPUs orchestrated as a single entity, which is essential for training advanced AI models.
Think SMART: How to Optimize AI Factory Inference Performance
The Think SMART framework enables enterprises to optimize AI factory inference by balancing accuracy, latency, and ROI, essential for handling complex AI models and diverse workloads effectively.
MindJourney enables AI to explore simulated 3D worlds to improve spatial interpretation
MindJourney empowers AI to navigate 3D environments by simulating movement and generating multiple perspectives, enhancing spatial reasoning capabilities beyond static images.
Applicability vs. job displacement: further notes on our recent research on AI and occupations
AI's applicability in occupations is primarily beneficial for knowledge work and communication tasks, such as writing and information gathering, rather than job displacement, which the study explicitly cautions against concluding.