Gemma 4 12B: A unified, encoder-free multimodal model
Gemma 4 12B introduces a novel unified architecture that eliminates multimodal encoders, allowing direct integration of audio and visual inputs into the LLM backbone for enhanced performance.
The ways we contain Claude across products
Anthropic has evolved its approach to deploying Claude, allowing greater access while implementing robust containment strategies to mitigate risks associated with model capabilities and user interactions. This shift reflects a growing confidence in the safety of AI systems as safeguards improve, balancing productivity gains against potential hazards.
KVarN: Native vLLM KV-cache quantization back end by Huawei
KVarN enhances KV-cache capacity by 3-5x and boosts throughput by ~1.3x while maintaining FP16-level accuracy, making it ideal for long-context workloads and concurrent requests.
Journey to JPEG XL: open-source experiments shaped the future of image coding
JPEG XL emerged from a decade of open-source experimentation, enhancing image coding through innovations like Butteraugli and Brotli, which improved compression efficiency and visual fidelity.
thunderbolt-ibverbs: We have InfiniBand at home
Experimental RDMA-over-USB4 enables 128GB Strix Halo mini PCs to communicate at ~95 Gb/s bidirectional speeds, facilitating tensor-parallel inference and FSDP workloads without enterprise gear.
MiniMax dropped a new attention architecture.
MiniMax Sparse Attention (MSA) achieves 1M tokens scaling by innovating memory access patterns, enhancing efficiency without sacrificing recall through a unique "KV outer gather Q" method.
Show HN: Mnemo – local-first AI memory layer for any LLM (Rust, SQLite, petgraph)
mnemo is a local-first AI memory layer that enables persistent knowledge management for LLMs, extracting entities and relationships to enhance future interactions without relying on cloud services.
On-policy distillation: one of the hottest terms on PapersWithCode
On-policy distillation (OPD) is a pivotal technique in AI, enhancing models like Qwen 3.6 and 3.7, GLM-5.1, and DeepSeek-V4 by refining error correction through targeted hint tokens.
KVarN: Variance-Normalized KV-Cache Quantization
KVarN introduces a novel KV-Cache quantization method that utilizes Hadamard rotations and variance-normalization on both K and V matrices, achieving 3-4x compression with minimal accuracy loss (0-1%) in decode-heavy tasks.
Direct Preference Optimization Beyond Chatbots
Direct Preference Optimization (DPO) extends beyond traditional chatbots, enabling more nuanced interactions and personalized user experiences in AI applications.
NVIDIA Research Unlocks Advanced Grasping, Smarter Autonomous Driving and Agent Training at Scale
NVIDIA's research introduces GraspGen-X, a foundation model for zero-shot grasping, enabling robots to adapt to new grippers and objects without retraining, thus enhancing versatility in robotic applications.
NVIDIA Enables the Next Era Of Physical AI Research With Agent Skills For Autonomous Vehicles, Robotics And Vision AI
NVIDIA's new physical AI agent skills, powered by NVIDIA Cosmos 3, streamline the development of autonomous vehicles, robotics, and vision AI, enhancing data generation and simulation workflows.