The last six months in LLMs have seen significant advancements, particularly marked by the November 2025 inflection point, which catalyzed improvements in coding agents and model performance.
Project Glasswing: what Mythos showed us
Mythos Preview represents a significant advancement in vulnerability detection, enabling the construction of exploit chains and proof generation, which enhances its ability to identify and validate security flaws effectively.
Anthropic acquires Stainless
Anthropic's acquisition of Stainless enhances its capabilities in SDK generation and agent connectivity, crucial for the evolving landscape of AI where agents must act effectively within their environments.
Cursor Introduces Composer 2.5
Composer 2.5 significantly enhances intelligence and collaboration, outperforming its predecessor by effectively managing long tasks and complex instructions, thanks to improved training methods and targeted reinforcement learning with textual feedback.
Reviving PapersWithCode (by Hugging Face)
Hugging Face is revitalizing PapersWithCode, utilizing AI to parse high-impact papers and generate leaderboards, ensuring the platform remains a vital resource for the ML community.
Gemini 3.5: frontier intelligence with action
Gemini 3.5 Flash introduces advanced agentic capabilities, enabling rapid execution of complex workflows and outperforming previous models in coding benchmarks, achieving 76.2% on Terminal-Bench 2.1 and 4x faster output than competitors.
Agora-1: The Multi-Agent World Model
Agora-1 introduces a multi-agent world model that allows up to four participants, human or AI, to interact in a shared, real-time simulation, enhancing experiences across various fields like gaming and robotics.
Vera Arrives: NVIDIA’s First CPU Built for Agents Lands at Top AI Labs
NVIDIA's Vera CPU is designed specifically for AI agents, marking a significant advancement in computational capabilities for AI applications in top labs.
Google changes its search box
Google's I/O 2026 introduces AI agents and a reimagined Search box, enhancing user interaction by allowing queries through natural language and providing intelligent suggestions, marking the most significant upgrade in over 25 years.
Cutting inference cold starts by 40x with LP, FUSE, C/R, and CUDA-checkpoint
Modal's engineering innovations have reduced inference cold start times by 40x, enabling rapid scaling of AI applications through techniques like cloud buffers, custom filesystems, and checkpoint/restore methods for both CPU and GPU memory.
Sub-JEPA: a simple fix to LeCun group's LeWorldModel that consistently improves performance
Sub-JEPA enhances LeCun's LeWorldModel by applying Gaussian regularization within multiple frozen random orthogonal subspaces, improving performance on low-dimensional tasks without introducing new hyperparameters.
KV Sharing, MHC, and Compressed Attention
Recent LLM architectures like Gemma 4 and DeepSeek V4 are innovating with techniques such as KV sharing and compressed attention to enhance long-context efficiency, significantly reducing memory and computational costs.
NVIDIA CEO Jensen Huang at Dell Technologies World: “Demand Is Going Parabolic, Utterly Parabolic”
NVIDIA's CEO Jensen Huang highlighted a dramatic surge in demand for AI technologies, with enterprise AI transitioning from pilot projects to large-scale deployments, driven by advancements in productivity and computation capabilities.
Introducing the Ettin Reranker Family
The Ettin Reranker Family introduces advanced models for text ranking, enhancing retrieval tasks with improved accuracy and efficiency, as seen in the cross-encoder/ettin-reranker-150m-v1 model.
Mistral AI Acquires Emmi AI to Create the Leading AI Stack
Mistral AI's acquisition of Emmi AI aims to create a leading AI stack for Industrial Engineering, leveraging Emmi's expertise in Physics AI to enhance industrial simulation and workflows across sectors like energy and aerospace.