ML Times

ML Times

R-Zero: Self-Evolving Reasoning LLM from Zero Data

R-Zero introduces a fully autonomous framework for training Large Language Models (LLMs) that generates its own data, eliminating reliance on human-curated tasks and labels, thus paving the way for super-intelligence.

Defeating Nondeterminism in LLM Inference

Nondeterminism in LLM inference arises from floating-point non-associativity and batch size variability, which can lead to different outputs even with identical inputs, as demonstrated by the varying completions generated by models like Qwen-3.

Implementation and ablation study of the Hierarchical Reasoning Model (HRM): what really drives performance?

The Hierarchical Reasoning Model (HRM) excels in performance primarily through outer-loop refinement with more segments, rather than its architecture, indicating a focus on training methodology over structural complexity.

LLMs play a cooperative card game, coordination without communication

LLMs struggle with cooperative gameplay, particularly in the card game The Crew, where smaller models fail to understand their roles, often prioritizing individual success over team strategy.

News: Arm announces next Generation core family called Arm Lumex

Arm's Lumex platform introduces C1 CPUs with Scalable Matrix Extensions 2 (SME2) and the Mali G1-Ultra GPU, designed specifically for AI applications in next-gen PCs and smartphones.

UGMM-NN: Univariate Gaussian Mixture Model Neural Network

The uGMM-NN introduces a novel neural architecture that embeds probabilistic reasoning into deep networks, allowing each node to parameterize activations as a univariate Gaussian mixture with learnable parameters.

Interesting PEZY-SC4s

PEZY-SC4S aims for highly efficient FP64 compute by utilizing a smaller die and lower power draw, achieving a projected ~91 GF/W performance, significantly outperforming Nvidia's H200 and approaching AMD's MI300A.

Graphrag pipeline that runs entirely locally with ollama and has full source attribution

VeritasGraph is a locally-run Graph RAG pipeline utilizing Ollama with Llama 3.1, designed for private use and ensuring full source attribution for generated content.

NVIDIA Blackwell Ultra crushes MLPerf

NVIDIA's Blackwell Ultra achieved 5× throughput on DeepSeek-R1 and set records on Llama 3.1 and Whisper, showcasing innovative techniques like FP8 KV-cache and disaggregated serving.

Stability AI Introduces Stable Audio 2.5, the First Audio Model Built for Enterprise Sound Production at Scale

Stable Audio 2.5 is the first audio generation model tailored for enterprise-grade sound production, enabling brands to create distinct audio identities across various channels, enhancing memorability by up to eight times.

Jupyter Agents: training LLMs to reason with notebooks

Jupyter Agents empower LLMs to execute code within Jupyter Notebooks, enhancing their ability to tackle complex data analysis tasks through direct code execution and reasoning.

NVIDIA Blackwell Ultra Sets the Bar in New MLPerf Inference Benchmark

The NVIDIA GB300 NVL72 rack-scale system achieves record-breaking throughput on the new reasoning inference benchmark, outperforming previous models by up to 1.4x in MLPerf Inference v5.1.

NVIDIA Partners With AI Infrastructure Ecosystem to Unveil Reference Design for Giga-Scale AI Factories

NVIDIA's new reference design aims to revolutionize data centers into integrated AI factories, enhancing energy efficiency and performance through collaboration with industry partners like Jacobs and Siemens Energy.

Breaking the networking wall in AI infrastructure

MOSAIC technology aims to resolve the power, reliability, and reach trade-off in AI infrastructure by utilizing a novel optical link design that combines low power consumption with high reliability and long distances, achieving up to 50 meters of reach.