ML Times
Jun 8, 2025
The last six months in LLMs, illustrated by pelicans on bicycles
The last six months in LLMs have seen over 30 significant model releases, highlighting the rapid evolution of the field, with notable advancements in local model capabilities, such as Meta's Llama 3.3 70B, which can run on consumer hardware and performs comparably to larger models.
Large Reasoning Models (LRMs) exhibit a collapse in accuracy when faced with complex problems, revealing their limitations in reasoning despite advanced capabilities.
Large Reasoning Models (LRMs) exhibit improved performance on reasoning tasks but face a collapse in accuracy beyond certain complexities, revealing critical limitations in their reasoning capabilities.
Ryan Williams' breakthrough demonstrates that all algorithms can be simulated with significantly less memory than previously thought, specifically showing that DTIME((t(n))) is a subset of DSPACE((\sqrt{t(n)\log t(n)})), a vast improvement over the classic 1977 result.
Log-Linear Attention enhances Mamba2 by allowing a dynamic state that grows over time, significantly improving long-range performance in sequence processing.
Apple's recent paper presents a compelling critique of LLMs, revealing their inability to reliably perform reasoning tasks, such as the classic Tower of Hanoi, even when provided with solution algorithms.
LLMs can write code, but they require intense supervision from skilled engineers, as demonstrated by the successful creation of a standards-compliant HTTP/2 server, which involved meticulous context management and workflow micromanagement.
Neural Differential-Algebraic Equations (DAEs) enable the integration of hard constraints into machine learning models, allowing for precise control over system behavior during training and simulation, as demonstrated in the manuscript Semi-Explicit Neural DAEs.
Gemini Diffusion excels in reasoning tasks and operates at remarkable speed, indicating its potential for transformative applications in AI.
The Apple paper titled The Illusion of Thinking reveals that reasoning models struggle with complex tasks, often giving up rather than attempting to solve them, indicating a potential inherent compute scaling limit.
This study reveals that reinforcement learning can enhance large language models (LLMs) to not only predict but also explain human decision-making processes, offering a dual-functionality that is often lacking in traditional models.