# Dec 1, 2025

- **A New AI Winter Is Coming**  
  - **LLMs are fundamentally flawed**, as they generate plausible but often incorrect outputs, leading to a high failure rate in practical applications, with estimates suggesting **60% to 95%** of results may be erroneous.

- **Program-of-Thought Prompting Outperforms Chain-of-Thought by 15% (2022)**  
  - **Program of Thoughts (PoT)** enhances language models by **separating reasoning from computation**, using Codex to articulate reasoning as a program and an external computer for calculations.

- **I Tested the M5 iPad Pro's Neural-Accelerated AI, and the Hype Is Real**  
  - The **M5 iPad Pro** demonstrates a **4.4× improvement** in time to first token (TTFT) for local AI processing, significantly exceeding Apple's claims, particularly for long prompts of 10,000 and 16,000 tokens.

- **Every Token Counts: Generalizing 16M Ultra-Long Context in Large Language Models**  
  - This research introduces **Hierarchical Sparse Attention (HSA)**, a novel mechanism that enables efficient ultra-long context modeling by ensuring **sparsity**, **random-access flexibility**, and **length generalization**.

- **🤗Transformers v5: Simple model definitions powering the AI ecosystem**  
  - **Transformers v5** introduces a **modular design** that simplifies model integration, allowing for **faster contributions** and a cleaner codebase, enhancing the overall user experience in the AI ecosystem.

- **At NeurIPS, NVIDIA Advances Open Model Development for Digital and Physical AI**  
  - **NVIDIA introduces the DRIVE Alpamayo-R1**, the first open industry-scale reasoning vision language action model for autonomous driving, enhancing vehicle decision-making in complex scenarios through advanced AI reasoning techniques.

- **Pose-free 3D Gaussian splatting via shape-ray estimation**  
  - **SHARE** introduces a **pose-free** framework for **3D Gaussian splatting**, enabling efficient rendering without the need for precise camera poses, thus addressing common challenges in real-world applications.

- **LFM2 Technical Report**  
  - **LFM2** is a family of **Liquid Foundation Models** optimized for **on-device deployment**, achieving up to **2x faster prefill and decode** on CPUs through a compact hybrid backbone of gated short convolutions and grouped query attention blocks.

- **Outcome-based learning vs vector search: 100% vs 3.3% accuracy on adversarial queries (p=0.001)**  
  - **Outcome-based learning achieved a remarkable 100% accuracy on adversarial queries**, significantly outperforming vector search's 3.3% accuracy, highlighting the potential of integrating outcome effectiveness into retrieval systems.

- **Multi-chain Graph Refinement and Selection for Reliable Reasoning in Large Language Models**  
  - The **Multi-chain Graph Refinement & Selection (MGRS)** framework enhances reasoning in Large Language Models (LLMs) by generating diverse trajectories and refining responses through a composite verification strategy, addressing limitations of existing methods like Tree-of-Thought and Graph-of-Thought.

- **TIM-PRM: Verifying multimodal reasoning with Tool-Integrated PRM**  
  - **TIM-PRM** revolutionizes verification in **Multimodal Large Language Models (MLLMs)** by transforming it into an active, tool-augmented investigation, addressing vulnerabilities like visual hallucinations and logical inconsistencies that traditional methods fail to resolve.

- **ORION: Teaching Language Models to Reason Efficiently in the Language of Thought**  
  - **ORION** introduces a framework that leverages the **Language of Thought Hypothesis** to train models for **ultra-compressed reasoning**, significantly enhancing efficiency in complex problem-solving.

- **Efficient MoE Pre-training at Scale on 1K AMD GPUs with TorchTitan**  
  - **Efficient training of Mixture-of-Experts (MoE) models** like DeepSeek-V3 and Llama 4-Scout was achieved using **TorchTitan** and **Primus-Turbo** on **1,024 AMD MI325X GPUs**, demonstrating a **2.77× speed-up** and **96% scaling efficiency**.
