Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
Needle distills Gemini 3.1 into a 26m parameter Simple Attention Network, enabling local finetuning on personal devices, achieving 6000 toks/sec prefill and 1200 decode speed in production.
A History of IDEs at Google
Google's IDE landscape evolved from fragmentation to a unified platform with the introduction of Cider V, which now supports 80% of development in the main codebase, enhancing productivity through better integrations and AI features.
The US is winning the AI race where it matters most: commercialization
The US leads in AI commercialization, leveraging cloud infrastructure, data, and strategic capital, significantly outpacing competitors like China in revenue and adoption since the launch of DeepSeek R1 in January 2025.
An idiot's guide to lead optimisation for proteins
Lead optimisation is a critical phase in drug design, where existing molecules are refined to enhance their efficacy, leveraging machine learning to propose and test modifications efficiently, as exemplified by the Cradle-1 pipeline.
MinT: Managed Infrastructure for Training and Serving Millions of LLMs
MinT is a managed infrastructure system that optimizes Low-Rank Adaptation (LoRA) for training and serving millions of large language models (LLMs) by efficiently handling adapter revisions without the need for full model checkpoints.
Human-level performance via ML was not proven impossible with complexity theory
The Ingenia Theorem proposed by Van Rooij et al. claims that achieving AGI via ML is impossible, but this assertion is fundamentally flawed due to the lack of a precise definition for "human-level classifier."
Hermes Unlocks Self-Improving AI Agents, Powered by NVIDIA RTX PCs and DGX Spark
Hermes introduces self-improving AI agents that leverage NVIDIA RTX PCs and DGX Spark, enabling continuous learning and adaptation through user feedback and task complexity.
Marco Polo: Finding a friend with only distance and motion
The article explores the challenge of locating devices in crowded spaces using range-only relative localization, leveraging ultra-wideband (UWB) technology and inertial measurement units (IMUs) to track movement and distance.
Elastic Attention Cores for Scalable Vision Transformers
Elastic Attention Cores introduce a core-periphery block-sparse attention structure for Vision Transformers, reducing computational costs from N² to 2NC + C² for C core tokens, enhancing scalability at higher resolutions.
LLMs are breaking 20 year old system design
LLMs challenge the 20-year-old web architecture by introducing long-running, stateful processes that require a new routing primitive, as traditional models assume stateless interactions with databases.
EditLens: Quantifying the extent of AI editing in text (2025)
AI editing is prevalent, with a significant number of queries to language models focused on editing rather than generating new content, revealing a need for tools like EditLens to quantify this phenomenon.
Trained transformer-based chess models to play like humans (including thinking time)
Trained transformer-based chess models achieve superior accuracy compared to MAIA-2 and are competitive with MAIA-3, utilizing nearly 1B games from Lichess across various rating buckets from ~800 to 2500+ with only 9MM parameters.
Learning, Fast and Slow: Towards LLMs That Adapt Continually
Fast-Slow Training (FST) enables large language models (LLMs) to adapt continually by utilizing fast weights for task-specific learning while maintaining slow weights for general reasoning, enhancing efficiency and performance.
Continual Harness: Online Adaptation for Self-Improving Foundation Agents
Continual Harness automates the iterative refinement process for self-improving foundation agents, enhancing their performance through model-harness co-learning, as demonstrated by the Gemini Plays Pokémon project.
Scenema Audio: Zero-shot expressive voice cloning and speech generation
Scenema Audio enables zero-shot expressive voice cloning, allowing any voice to convey various emotions without prior recordings, enhancing creative workflows in video production.