ML Times
Apr 11, 2025
Yann LeCun Auto-Regressive LLMs are Doomed
Yann LeCun argues that auto-regressive LLMs are fundamentally flawed, suggesting that their reliance on sequential data processing limits their potential for true understanding and reasoning.
2025 AI Index Report
The 2025 AI Index Report reveals significant advancements in AI research, highlighting a 50% increase in published papers compared to the previous year, indicating a robust growth in the field.
Controlling Language and Diffusion Models by Transporting Activations
Activation Transport (AcT) is a novel framework that enables fine-grained control over generative models' outputs with minimal computational overhead, addressing challenges in reliability and safety without altering model parameters.
Rust CUDA Project
The Rust CUDA Project aims to establish Rust as a tier-1 language for high-performance GPU computing, leveraging the CUDA Toolkit to compile Rust into optimized PTX code and integrate with existing CUDA libraries.
B200 vs H100 Benchmarks: Early Tests Show Up to 57% Faster Training Throughput & Self-Hosting Cost Analysis
B200 GPUs demonstrate up to 57% higher training throughput than H100s in computer vision tasks, indicating a significant performance leap for model training workloads.
Levitating Bugs with Sound Could Transform Scientific Photography
Researchers have developed a method using acoustic levitation to capture detailed photographs of insect specimens without physical damage, enabling automated imaging from multiple angles.
A16Z: AI Avatars
AI avatars are evolving beyond static images to create dynamic, talking characters that combine realistic facial expressions, body language, and voice, enhancing user engagement in content creation and marketing.
Agency vs. Control vs. Reliability in Agent Design
The ACR tradeoff highlights that while high-agency AI agents can perform complex tasks, they often lack the reliability needed for effective customer support, necessitating a balance between autonomy and control.
Show HN: Lunon – Instant model switching across LLMs
Lunon enables seamless LLM model swapping and deployment in milliseconds, allowing users to integrate new models on the same day they become available, thus enhancing operational efficiency.
We built an OS-like runtime for LLMs — curious if anyone else is doing something similar?
The team has developed an AI-native runtime that can snapshot-load LLMs (13B–65B) in 2–5 seconds, enabling the dynamic execution of 50+ models per GPU without constant memory residency.
Debug-gym: an environment for AI coding tools to learn how to debug code like programmers
Debug-gym is a novel environment designed to enhance AI coding tools' debugging capabilities by allowing them to interactively seek information and propose fixes, thereby mimicking human debugging processes.
Sub-2s cold starts for 13B+ LLMs + 50+ models per GPU — curious how others are tackling orchestration?
Snapshot-loading LLMs (13B–65B) in under 2–5 seconds enables dynamic execution of 50+ models per GPU, enhancing efficiency without constant memory residency.
A slop forensics toolkit for LLMs: computing over-represented lexical profiles and inferring similarity trees
I built a biomedical GNN + LLM pipeline (XplainMD) for explainable multi-link prediction
SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
SpecReason accelerates LRM inference by utilizing a lightweight model for intermediate reasoning, achieving 1.5-2.5× speedup while enhancing accuracy by 1.0-9.9%.