ML Times

Jun 4, 2026

Gemma 4 12B: A unified, encoder-free multimodal model

The ways we contain Claude across products

KVarN: Native vLLM KV-cache quantization back end by Huawei

Journey to JPEG XL: open-source experiments shaped the future of image coding

thunderbolt-ibverbs: We have InfiniBand at home

MiniMax dropped a new attention architecture.

Show HN: Mnemo – local-first AI memory layer for any LLM (Rust, SQLite, petgraph)

On-policy distillation: one of the hottest terms on PapersWithCode

KVarN: Variance-Normalized KV-Cache Quantization

Direct Preference Optimization Beyond Chatbots

NVIDIA Research Unlocks Advanced Grasping, Smarter Autonomous Driving and Agent Training at Scale

NVIDIA Enables the Next Era Of Physical AI Research With Agent Skills For Autonomous Vehicles, Robotics And Vision AI