ML Times

Flash-MoE: Running a 397B Parameter Model on a Laptop

Flash-MoE enables running a 397 billion parameter Mixture-of-Experts model on a MacBook Pro at 4.4+ tokens/second, utilizing a pure C/Metal inference engine without Python or frameworks.


Has industry effectively killed off academic machine learning research in 2026?

Industry dominance in machine learning research has surged, leveraging vast resources and talent, leaving academia to focus on niche topics and theoretical scenarios that lack practical application.


The quadratic problem nobody fixed

Regex engines universally exhibit O(n²) complexity when finding all matches, despite claims of linear performance; this issue has persisted since the 1970s, affecting all languages and implementations.


Dataframe 1.0.0.0

DataFrame 1.0.0.0 introduces a typed API that ensures compile-time schema validation, enhancing the transition between exploratory and pipeline work, thanks to community feedback from contributors like maxigit and mcoady.


I Reverse-Engineered the TiinyAI Pocket Lab from Marketing Photos

TiinyAI's Pocket Lab, marketed as a "pocket-sized AI supercomputer," is built on a CIX P1 SoC and a dual-die VeriSilicon VIP9400 NPU, but its performance is severely limited by a split memory architecture that hinders inference speed. The device claims to run a 120-billion-parameter model at 20 tokens per second, yet the actual active parameters are only 5.1 billion, significantly reducing its effective performance.


Intuitions for Transformer Circuits

Mechanistic Interpretability (MI) is crucial for understanding transformer models, as it allows us to reverse-engineer their behavior and ensure alignment with human values, addressing the risks posed by AI systems.


AI Risks "Hypernormal" Science

Scaling AI does not guarantee paradigm shifts; true innovation requires AI to generate new conceptual frameworks rather than merely enhancing existing predictive capabilities.


The $\mathbf{Y}$-Combinator for LLMs: Solving Long-Context Rot with $λ$-Calculus

$λ$-RLM revolutionizes long-context reasoning in LLMs by utilizing typed functional runtime from $λ$-calculus, ensuring structured execution and enhanced reliability over traditional RLMs.


Modeling online discourse escalation as a state machine (dataset + labeling approach)

The proposed framework models online discourse escalation as a state machine, identifying seven distinct states from Neutral to Threats of violence, with each comment labeled according to its local state while the thread evolves globally.


How Autonomous AI Agents Become Secure by Design With NVIDIA OpenShell

NVIDIA OpenShell enhances the security of autonomous AI agents by implementing a trusted infrastructure policy layer, ensuring that security measures are applied at the system level rather than relying on the agents themselves.


Detection Is Cheap, Routing Is Learned: Why Refusal-Based Alignment Evaluation Fails (arXiv 2603.18280)

Refusal-based benchmarks fail to capture the essential learned routing mechanisms in alignment evaluation, as demonstrated through a study on political censorship in Chinese-origin LLMs, revealing that these mechanisms are lab-specific and often invisible to traditional methods.


Solving the "Liquid-Solid Interface" Problem: 116 High-Fidelity Datasets of Coastal Physics (Waves, Saturated Sand, Light Transport)


Training a classifier entirely in SQL (no iterative optimization)


The Residual Stream Is All You Need: On the Redundancy of the KV Cache in Transformer Inference

The key-value (KV) cache in transformer inference is proven to be redundant, as keys and values can be perfectly reconstructed from the residual stream, yielding zero reconstruction error across various models, including those with 135M to 4B parameters.


Semantic Token Clustering for Efficient Uncertainty Quantification in Large Language Models

Semantic Token Clustering (STC) enhances uncertainty quantification in large language models (LLMs) by grouping tokens into semantically consistent clusters, allowing for efficient identification of unreliable outputs without the need for repeated sampling or auxiliary models.