ML Times
Oct 31, 2025
Reasoning Models Reason Well, Until They Don't
Large reasoning models (LRMs) demonstrate impressive performance on reasoning tasks but fail catastrophically when faced with complex problems, revealing a critical limitation in their capabilities.Kimi Linear: An Expressive, Efficient Attention Architecture
Kimi Linear introduces a hybrid linear attention architecture that significantly enhances performance and efficiency, achieving 51.0 on MMLU-Pro and 84.3 on RULER with a 3.98x speedup over traditional methods.Sustainable memristors from shiitake mycelium for high-frequency bioelectronics
Shiitake mycelium serves as a sustainable substrate for memristors, achieving operational frequencies up to 5.85 kHz with an accuracy of 90 ± 1%, showcasing its potential in bioelectronics and neuromorphic computing.FastJAM: a Fast Joint Alignment Model for Images (NeurIPS 2025)
FastJAM introduces a lightweight graph-based framework for joint image alignment, achieving results in seconds compared to the hours required by previous methods, utilizing sparse keypoints and graph neural networks (GNNs) for enhanced efficiency.Exceptional Measurement of Chirality
Breakthrough research enhances the measurement of flexible left- and right-handed molecules, known as chiral molecules, using a novel approach to Vibrational Circular Dichromism (VCD) spectroscopy, enabling real-time monitoring of biochemical processes and high-throughput pharmaceutical screening.We found LRMs look great…until the problems get harder (AACL 2025)
Large Reasoning Models (LRMs) excel in simpler tasks but exhibit a sharp decline in performance as reasoning complexity increases, revealing a critical threshold beyond which they struggle.triton_bwd: Enabling Backpropagation for the OpenAI Triton language
Triton language enables fast GPU kernel coding in Python, but lacks straightforward backpropagation for custom operations; a new library, triton_bwd, introduces automatic differentiation for Triton kernels.Update: Added Full Drift Benchmark Report (PKBoost vs LightGBM vs XGBoost — 16 Scenarios)
PKBoost outperforms LightGBM and XGBoost by +50-60% in PR AUC gains, demonstrating superior performance across 16 drift scenarios, with an average PR-AUC of 0.8509 and minimal degradation of only 2.82%.Layer-0 heads that pre-bias hedging over facts in GPT-2 (replicated in Mistral-7B) — code + DOI
Layer-0 heads in GPT-2 significantly downweight factual continuations while boosting hedging tokens, with zeroing specific heads improving logit-difference by +0.40–0.85 and enhancing calibration metrics (ECE 0.122→0.091, Brier 0.033→0.024).How to benchmark open-ended, real-world goal achievement by computer-using LLMs?
Benchmarking open-ended, real-world goal achievement by LLMs reveals that while agents excel in verbal tasks, they struggle with real-world actions, often leading to inflated self-assessments of their performance.Korea Joins AI Industrial Revolution: NVIDIA CEO Jensen Huang Unveils Historic Partnership at APEC Summit
South Korea's AI initiative involves deploying 260,000 NVIDIA GPUs across various sectors, marking a significant leap towards a sovereign AI infrastructure that aims to transform industries like manufacturing and robotics.A New Species of Artificial Intelligence: KMS-Stabilized Reasoning with Harmonic Algebra
KMS-stabilized reasoning with harmonic algebra proposes a theoretical framework for AI that surpasses classical limits, enabling continuous processing and formal stability guarantees, which could lead to exponential speedups in specific problem classes.ExpertFlow: Adaptive Expert Scheduling and Memory Coordination for Efficient MoE Inference
ExpertFlow enhances Mixture-of-Experts (MoE) inference by integrating adaptive expert prefetching and cache-aware routing, significantly reducing memory demand and computational overhead during model execution.Stop Wasting Your Tokens: Towards Efficient Runtime Multi-Agent Systems
SupervisorAgent introduces a lightweight framework for runtime supervision in Multi-Agent Systems (MAS), enhancing efficiency by reducing token consumption by an average of 29.45% while maintaining success rates.