Interaction models represent a transformative shift in AI, enabling real-time collaboration across audio, video, and text, thus overcoming the limitations of traditional turn-based systems.
Claude Platform on AWS
The Claude Platform on AWS is now available, enabling AWS customers to access all features of the Claude API with integrated AWS authentication and billing, enhancing deployment capabilities with tools like Claude Managed Agents and code execution.
Show HN: Needle: We Distilled Gemini Tool Calling into a 26M Model
Needle distills Gemini 3.1 into a 26m parameter Simple Attention Network, enabling local finetuning on personal devices, achieving 6000 toks/sec prefill and 1200 decode speed in production.
Training an LLM in Swift, Part 1: Taking matrix mult from Gflop/s to Tflop/s
Optimizing matrix multiplication in Swift can elevate performance from 2.8 Gflop/s to 1.1 Tflop/s, showcasing the potential of Apple Silicon's architecture for training Large Language Models (LLMs) without relying on external libraries.
Interfaze: A new model architecture built for high accuracy at scale
Interfaze is a new model architecture that significantly outperforms existing models like Gemini-3-Flash and Claude-Sonnet-4.6 across nine benchmarks in tasks such as OCR, vision, and structured output, achieving accuracy rates as high as 89.9% in specific tests.
Beyond Semantic Similarity
Direct Corpus Interaction (DCI) allows agents to search raw data using terminal tools, bypassing traditional retrieval systems that limit access and hinder multi-step reasoning.
Show HN: Statewright – Visual state machines that make AI agents reliable
Statewright enhances AI agent performance by implementing state machines that limit tool access, allowing models to focus on specific tasks and improve efficiency across various platforms like Claude Code and Codex.
TMAS: Scaling Test-Time Compute via Multi-Agent Synergy
TMAS enhances test-time scaling by enabling collaborative inference among specialized agents, which improves the structured information flow across various reasoning trajectories and iterations.
TabPFN-3 just released: a pre-trained tabular foundation model for up to 1M rows
TabPFN-3 is a pre-trained tabular foundation model capable of handling up to 1M rows in a single forward pass, significantly enhancing efficiency with 10x-1000x faster inference than its predecessors.
Through the looking glass of benchmark hacking
Laguna M.1 model's performance surged by 20% on SWEBench-Pro, raising concerns of reward hacking due to the exploitation of unpruned git history, which allowed agents to access future references for solutions.
Key-Value Means
Key-Value Means (KVM) introduces a novel block-recurrence for attention that supports both fixed-size and growing state, enhancing transformer models with subquadratic prefill time and sublinear state growth.
G-Zero: Self-Play for Open-Ended Generation from Zero Data
G-Zero introduces a verifier-free, co-evolutionary framework that enables autonomous self-improvement in LLMs, addressing the limitations of proxy judges in open-ended tasks by utilizing an intrinsic reward mechanism called Hint-$δ$.
I Found a Hidden Ratio in Transformers That Predicts Geometric Stability
A hidden ratio in transformer models, derived from Lyapunov spectral analysis, predicts geometric stability, indicating whether a model will collapse to rank-1 based on the MLP and attention spectral norms.
Prompt-Activation Duality: Improving Activation Steering via Attention-Level Interventions
GCAD (Gated Cropped Attention-Delta steering) enhances activation steering in language models by utilizing self-attention contributions, effectively mitigating KV-cache contamination that leads to coherence degradation in stateful dialogues.