ML Times
Jun 7, 2025
Daily
Sandia National Labs has activated its SpiNNaker 2 supercomputer, a brain-inspired system that operates without traditional storage, mimicking 150 to 180 million neurons to enhance computational efficiency for national security applications.
LLMs can be transformed into locally linear systems that accurately reconstruct next-token outputs without altering model weights, enhancing interpretability through linear algebra techniques.
The highly efficient matrix transpose kernel implemented in Mojo achieves a remarkable bandwidth of 2775.49 GB/s, demonstrating that Mojo can match CUDA performance on the same task, with optimizations leading to significant speed improvements.
Log-linear attention enhances the efficiency of sequence modeling by replacing fixed-size hidden states with a logarithmically growing set of hidden states, achieving a compute cost that is log-linear in sequence length.
Open source LLMs are increasingly outperforming closed source models like GPT-4o-mini and Gemini 2.5 Flash for common tasks, offering significant cost savings and improved performance, especially in batch processing scenarios.
Cursor has secured $900 million in Series C funding, elevating its valuation to $9.9 billion, with backing from prominent investors like Thrive, Accel, and Andreessen Horowitz.
Log-Linear Attention enhances Mamba2 by allowing a dynamic state that grows over time, significantly improving long-range performance in sequence processing.
Large Reasoning Models (LRMs) exhibit a collapse in accuracy when faced with complex problems, revealing their limitations in reasoning despite advanced capabilities.
The latest Apple paper reveals that LLMs (Large Language Models) and LRMs (Language Reasoning Models) exhibit limited true reasoning capabilities, particularly struggling with complex tasks that require human-like reasoning.
Yet Another Quantization Algorithm (YAQA) significantly enhances model output preservation post-quantization, achieving a KL reduction of over 30% compared to QTIP and outperforming Google's QAT model on Gemma 3.
Gemini Diffusion excels in reasoning tasks and operates at remarkable speed, indicating its potential for transformative applications in AI.
A 100M open source notebooklm speech model has been developed using two NVIDIA 4090 GPUs, showcasing significant advancements in speech processing capabilities.
Ryan Williams' breakthrough demonstrates that all algorithms can be simulated with significantly less memory than previously thought, specifically showing that DTIME((t(n))) is a subset of DSPACE((\sqrt{t(n)\log t(n)})), a vast improvement over the classic 1977 result.
Bifrost is an open-source LLM gateway built in Go, engineered for high-throughput and low-latency deployments, addressing the common infrastructure challenges faced when scaling LLMs in production.