Waymo's safety data reveals a significant reduction in crash rates, with a 92% decrease in serious injury or worse crashes compared to human drivers, showcasing the effectiveness of autonomous driving technology in enhancing road safety.
Nvidia greenboost: transparently extend GPU VRAM using system RAM/NVMe
nvidia_greenboost is a project with 34 commits and 1 branch, showcasing active development in optimizing NVIDIA GPU performance.
Show HN: Duplicate 3 layers in a 24B LLM, logical deduction .22→.76. No training
Duplicating specific layers in transformer models can enhance reasoning capabilities significantly, with a 17% boost in Qwen2.5-32B and logical deduction improving from 0.22 to 0.76 in Devstral-24B, achieved without training or weight changes.
EsoLang-Bench: Evaluating Genuine Reasoning in LLMs via Esoteric Languages
EsoLang-Bench introduces a new benchmark for evaluating LLMs using esoteric programming languages, revealing that models trained on scarce data (5,000 to 100,000x less than Python) struggle significantly with code generation tasks.
Machine Payments Protocol (MPP)
The Machine Payments Protocol (MPP) enables autonomous agents to transact seamlessly, eliminating the cumbersome steps of traditional payment systems, thus fostering a new era of agent-driven commerce.
[D] ICML rejects papers of reviewers who used LLMs despite agreeing not to
ICML has taken a bold stance by rejecting all papers from reviewers who utilized LLMs despite their prior agreement to avoid such tools, marking a significant shift in conference review practices.
Measuring progress toward AGI: A cognitive framework
Google DeepMind introduces a cognitive framework to measure progress toward Artificial General Intelligence (AGI), emphasizing the need for empirical tools to evaluate AI systems' cognitive capabilities.
Book: The Emerging Science of Machine Learning Benchmarks
Machine Learning benchmarks are crucial for evaluating model performance, guiding researchers in developing more effective algorithms and systems.
Pretraining Language Models via Neural Cellular Automata
Neural Cellular Automata (NCA) can enhance language model training by utilizing synthetic data, achieving a 6% perplexity gain and 1.6× faster convergence compared to traditional methods using natural language data.
Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
Karpathy's autoresearch achieved a 2.87% improvement in validation loss by utilizing 16 GPUs to run ~910 experiments in just 8 hours, demonstrating the power of parallel processing in machine learning research.
NanoGPT Slowrun: 10x Data Efficiency with Infinite Compute
NanoGPT Slowrun achieves 10x data efficiency, training an ensemble of 1.8B parameter models on 100M tokens, outperforming traditional models that require 1B tokens, thus allowing for enhanced model performance through compute scaling rather than data scaling.
[R] A Gradient Descent Misalignment — Causes Normalisation To Emerge
Gradient descent misalignment leads to a systematic divergence in activation updates, revealing that while parameters follow the steepest descent, activations do not, which may explain the efficacy of normalization techniques.
Gluon: Explicit Performance
Gluon serves as a Python frontend to Triton GPU ttg IR, offering developers explicit control over GPU kernel programming, which enhances performance by exposing compiler internals that were previously hidden.
[R] Extreme Sudoku as a constraint-satisfaction benchmark, solved natively without tools or CoT or solution backtracking
Extreme Sudoku serves as a constraint-satisfaction benchmark, revealing that leading LLMs achieve 0% accuracy, while the BDH architecture excels with 97.4% accuracy without external tools or backtracking.
Show HN: I built a P2P network where AI agents publish formally verified science
P2PCLAW is a peer-to-peer network enabling AI agents and researchers to share and validate scientific results through formal mathematical proof, ensuring that only rigorously verified claims are accepted.