# Mar 7, 2025

### Mistral OCR
- **Mistral OCR** is a cutting-edge Optical Character Recognition API that excels in understanding complex document elements, achieving **94.89% accuracy** in rigorous benchmark tests, outperforming competitors like Google Document AI and Azure OCR.

### Moscow-based global news network has infected Western AI tools
- A **NewsGuard audit** reveals that leading AI chatbots repeated **33%** of false narratives from the Kremlin-affiliated Pravda network, indicating a significant infiltration of Russian propaganda into Western AI tools.

### Ladder: Self-Improving LLMs Through Recursive Problem Decomposition
- **LADDER** (Learning through Autonomous Difficulty-Driven Example Recursion) empowers Large Language Models to enhance their problem-solving skills by autonomously generating simpler problem variants, achieving a remarkable accuracy increase from **1% to 82%** in mathematical integration tasks.

### Differentiable Logic Cellular Automata
- **Differentiable Logic Cellular Automata (DiffLogic CA)** merges Neural Cellular Automata (NCA) with Differentiable Logic Gate Networks, enabling the learning of local rules for complex patterns while maintaining a discrete state space, thus enhancing interpretability and efficiency in computation.

### Using GRPO to Beat o1, o3-mini and R1 at "Temporal Clue"
- **GRPO** outperformed models like **R1**, **o1**, and **o3-mini** on the **Temporal Clue** reasoning task, achieving results comparable to **Sonnet 3.7** while being over **100x cheaper** at inference time.

### Why I find diffusion models interesting?
- **Diffusion LLMs (dLLMs)** represent a paradigm shift in language model architecture, enabling simultaneous word generation rather than the traditional left-to-right token prediction, leading to **5-10x improvements in speed and efficiency**.

### AMD YOLO
- **AMD's MI300X** boxes are on the way, signaling a potential shift in the AI hardware landscape as the company aims to challenge NVIDIA's dominance with a **sovereign software stack** that integrates seamlessly with PyTorch.

### Show HN: Open-source, native audio turn detection model
- **Smart turn detection** is an open-source model designed to enhance conversational AI by improving the accuracy of when voice agents should respond, moving beyond traditional **voice activity detection (VAD)** methods that fail to consider linguistic and acoustic nuances.

### Polars Cloud: The Distributed Cloud Architecture to Run Polars Anywhere
- **Polars Cloud** aims to unify DataFrame processing with a **flexible API** that supports **query optimization** and **parallel execution**, enabling seamless remote data processing across various cloud platforms like AWS, Azure, and GCP.

### InstantStyle: Free Lunch Towards Style-Preserving in Text-to-Image Generation
- **InstantStyle** introduces a novel framework for **decoupling style and content** in text-to-image generation, leveraging **CLIP global features** to effectively mitigate content leakage through a simple subtraction method.

### [P] Training a Rust 1.5B Coder LM with Reinforcement Learning (GRPO)
- The **1.5B Coder LM** achieved a notable improvement, increasing build pass rates from **~60% to ~80%** and unit test success from **22% to 37%** after training with **GRPO** on **15k examples**.

### [R] Cognitive Behaviors That Enable Language Model Self-Improvement: Analyzing Verification, Backtracking, Subgoals, and Backward Chaining
- **Four cognitive behaviors**— **double-checking**, **seeking background knowledge**, **step-back reasoning**, and **heuristic relaxation**—enable language models to enhance their reasoning capabilities autonomously, yielding significant performance gains across various reasoning tasks.

### Reflection – AlphaGo / Gemini team building superintelligent coding agents
- **Reflection AI** aims to create **superintelligent autonomous systems** that can perform cognitive tasks, drawing inspiration from the groundbreaking success of AlphaGo in 2016, which showcased the potential of AI to surpass human capabilities in complex problem-solving.

### L$^2$M: Mutual Information Scaling Law for Long-Context Language Modeling
- The **L$^2$M** framework introduces a **bipartite mutual information scaling law** that uniquely addresses long-range dependencies in language modeling, diverging from traditional two-point mutual information approaches.

### Study: Large language models still lack general reasoning skills
- **Large language models (LLMs) exhibit significant deficiencies in general reasoning skills**, indicating that their capabilities are still limited compared to human cognitive functions.
