NeurIPS 2025 awarded seven papers for their significant contributions to machine learning, covering topics like diffusion models, self-supervised reinforcement learning, and attention mechanisms in large language models, showcasing the conference's commitment to advancing AI research.
State of AI: An Empirical 100T Token Study with OpenRouter
The State of AI report from OpenRouter highlights significant advancements in machine learning and natural language processing, showcasing how these technologies are reshaping industries and enhancing productivity. Read more here.
Gemini 3 Pro: the frontier of vision AI
Gemini 3 Pro is Google's most advanced multimodal model, excelling in document, spatial, screen, and video understanding, enabling complex visual reasoning and document processing.
CUDA-l2: Surpassing cuBLAS performance for matrix multiplication through RL
CUDA-L2 leverages reinforcement learning to optimize Half-precision General Matrix Multiply (HGEMM) CUDA kernels, achieving superior performance compared to established libraries like cuBLAS and torch.matmul across 1,000 configurations on the A100 GPU.
Addressing the “LLMs with thousands of tools” challenge
Anthropic's Tool Search feature aims to address the “too many tools in context” issue by enabling models to discover tools on-demand, rather than preloading thousands of definitions.
Reframing Impact
An impact measure could serve as the first safeguard against powerful agents with imperfect objectives, enabling effective control without needing to define the objective itself. This approach emphasizes the importance of creating a robust framework that can mitigate risks associated with advanced AI systems.
Visualizing emergent structure in the Dragon Hatchling (BDH)
The BDH architecture offers a novel approach to pathfinding by modeling neuron-to-neuron interactions on sparse graphs, utilizing Hebbian learning to adapt its circuits dynamically, which distinguishes it from traditional transformer models.
Mitigating Object and Action Hallucinations in Multimodal LLMs via Self-Augmented Contrastive Alignment
The SANTA framework effectively mitigates object and action hallucinations in multimodal LLMs by focusing on visual facts and reducing spurious correlations, enhancing the accuracy of video descriptions.
Tiny Recursive Models (TRMs) leverage recursion to perform extensive computations with fewer parameters, enhancing efficiency in hierarchical reasoning tasks.
Semantic Soft Bootstrapping: Long Context Reasoning in LLMs without Reinforcement Learning
Semantic Soft Bootstrapping (SSB) enhances long context reasoning in large language models (LLMs) by utilizing a self-distillation technique that allows the model to act as both teacher and student, improving its reasoning capabilities without the need for reinforcement learning.
Multi-LLM Collaboration for Medication Recommendation
Multi-LLM collaboration enhances medication recommendations by addressing the hallucination and inconsistency issues prevalent in individual large language models (LLMs), utilizing a Chemistry-inspired framework for improved reliability.
Arbitrage: Efficient Reasoning via Advantage-Aware Speculation
Arbitrage introduces a dynamic routing mechanism that enhances step-level speculative generation, significantly improving the efficiency of reasoning tasks in Large Language Models by reducing unnecessary rejections of semantically equivalent steps.
SEASON: Mitigating Temporal Hallucination in Video Large Language Models via Self-Diagnostic Contrastive Decoding
SEASON introduces a novel method, Self-Diagnostic Contrastive Decoding, to combat temporal hallucination in Video Large Language Models (VideoLLMs), enhancing both temporal and spatial accuracy without requiring additional training.
AdmTree: Compressing Lengthy Context with Adaptive Semantic Trees
AdmTree introduces a novel framework for adaptive context compression in Large Language Models (LLMs), focusing on high semantic fidelity while enhancing computational efficiency through a hierarchical structure.
TaskEval: Synthesised Evaluation for Foundation-Model Tasks
TaskEval introduces a novel synthesised evaluator for Foundation Model (FM) tasks, addressing the critical issue of hallucinations in applications by automating evaluation and integrating human feedback.