ML Times
Sep 30, 2025
Claude Sonnet 4.5 is hailed as the best coding model globally, showcasing significant advancements in reasoning, math, and complex agent-building capabilities, with a 61.4% performance on the OSWorld benchmark, up from 42.2% just four months prior.
Julia's ecosystem suffers from a high rate of serious correctness bugs, undermining its reliability for critical applications, as evidenced by numerous filed issues, including incorrect results in core functions and package interactions.
Cerebras Systems has successfully raised $1.1 billion in a Series G funding round, achieving a post-money valuation of $8.1 billion, with major backing from Fidelity Management and a consortium of top investors.
RL with Zero-Variance Prompts (RL-ZVP) innovatively extracts learning signals from prompts that traditionally yield uniform rewards, enhancing the reasoning capabilities of large language models.
InfLLM-V2 introduces a dense-sparse switchable attention framework that allows seamless adaptation from short to long sequences, enhancing the efficiency of large language models without excessive parameters.
Future-Guided Learning enhances time-series forecasting by utilizing a dynamic feedback mechanism that aligns a forecasting model with future data, significantly improving prediction accuracy.
AceSearcher introduces a cooperative self-play framework that enhances LLMs by training them to alternate between decomposing complex queries and integrating retrieved contexts, significantly improving reasoning capabilities.
SWAX, a hybrid architecture of sliding-window attention and xLSTM linear RNN layers, demonstrates that short window attention enhances long-term memory retention, contrary to expectations about larger windows.
NVIDIA's Newton physics engine and enhanced Isaac GR00T models facilitate accelerated robot learning through OpenUSD simulation workflows, allowing for the training of numerous robot instances simultaneously using both real and synthetic data.
NVIDIA's CUDA-X libraries are revolutionizing quantum computing by enabling accelerated computing to tackle critical challenges like error correction and circuit optimization, thus paving the way for practical quantum applications.
DiDi-Instruct introduces a novel training method for language generation, achieving 64x acceleration over traditional models by leveraging a pre-trained discrete diffusion language model (dLLM).
SemShareKV introduces a novel framework for KV cache sharing that leverages fuzzy token matching via locality-sensitive hashing (LSH), enhancing inference efficiency for semantically similar prompts.
Semantic Curriculum Preference Optimization (SCPO) is a groundbreaking framework that significantly reduces visual hallucinations in Multimodal Large Language Models (MLLMs) by employing a structured, progressive learning approach based on fine-grained semantic contrasts.
This paper introduces an evolutionary framework for training large language models (LLMs), utilizing sub-networks called experts that share structure but differ in parameters, enhancing learning efficiency through evolutionary operators like crossover and mutation.