Waymo's safety data reveals a significant reduction in crash rates, with a 92% decrease in serious injury or worse crashes compared to human drivers, showcasing the effectiveness of autonomous driving technology in enhancing road safety.
MacBook M5 Pro and Qwen3.5 = Local AI Security System
Qwen3.5-9B achieves a remarkable 93.8% pass rate on the HomeSec-Bench, only 4.1 points behind GPT-5.4, demonstrating the potential of local AI solutions on consumer hardware like the MacBook Pro M5.
EsoLang-Bench: Evaluating Genuine Reasoning in LLMs via Esoteric Languages
EsoLang-Bench introduces a new benchmark for evaluating LLMs using esoteric programming languages, revealing that models trained on scarce data (5,000 to 100,000x less than Python) struggle significantly with code generation tasks.
Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
Karpathy's autoresearch achieved a 2.87% improvement in validation loss by utilizing 16 GPUs to run ~910 experiments in just 8 hours, demonstrating the power of parallel processing in machine learning research.
Flash-KMeans: Fast and Memory-Efficient Exact K-Means
Flash-KMeans transforms the traditional $k$-means algorithm into an efficient online processing tool, addressing critical IO bottlenecks and contention issues in GPU implementations.
NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute
NanoGPT Slowrun achieves 10x data efficiency, training an ensemble of 1.8B parameter models on 100M tokens, outperforming traditional models that require 1B tokens, thus allowing for enhanced model performance through compute scaling rather than data scaling.
Attention Residuals
Attention Residuals (AttnRes) enhances Transformer architectures by allowing layers to selectively aggregate earlier representations through learned attention, addressing the dilution of contributions in standard residual connections.
Parallel Perl – autoparallelizing interpreter with JIT
Richard Jelinek proposes a paradigm shift in programming by allowing AI to write Perl, leveraging decades of experience in AI and Perl development to enhance automation and efficiency in coding tasks.
Launch HN: Canary (YC W26) – AI QA that understands your code
Canary is an AI QA tool that intelligently analyzes code changes in pull requests, generating and executing tests to ensure user workflows function correctly, thus addressing the gap in pre-merge testing.
[R] Doc-to-LoRA: Learning to Instantly Internalize Contexts from Sakana AI
Doc-to-LoRA (D2L) introduces a lightweight hypernetwork that enables instant context internalization for Large Language Models (LLMs), allowing for efficient document understanding and multi-step reasoning without the need for extensive retraining.
Medical AI gets 66% worse when you use automated labels for training, and the benchmark hides it! [R][P]
Medical AI performance declines by 66% when trained with automated labels, revealing a significant risk in relying on such methods for segmentation tasks in breast cancer imaging.
NumKong: 2'000 Mixed Precision Kernels for All
NumKong offers over 2,000 mixed-precision SIMD kernels for various programming languages, optimizing performance across architectures like RISC-V, Intel AMX, and Arm SME, while maintaining a compact size of under 5 MB.
[D] Tried MiniMax M2.7 impressive performance on real-world tasks
MiniMax M2.7 excels in handling complex tasks, demonstrating impressive capabilities in coding workflows, bug tracing, and multi-step document edits, as experienced through ZenMux deployment.
🤗Introducing SPEED-Bench: A Unified and Diverse Benchmark for Speculative Decoding
SPEED-Bench introduces a comprehensive benchmark for evaluating speculative decoding (SD) across diverse semantic domains and realistic serving conditions, addressing the limitations of existing benchmarks that often lack scale and diversity.
NVIDIA GTC 2026: Live Updates on What’s Next in AI
NVIDIA and Thinking Machines Lab have forged a multiyear strategic partnership to deploy one gigawatt of NVIDIA Vera Rubin systems, enhancing frontier model training capabilities.