# Mar 6, 2025

### Apple M3 Ultra
- **M3 Ultra** delivers **up to 2.6x the performance** of M1 Ultra, featuring a **32-core CPU**, **80-core GPU**, and support for **over 512GB of unified memory**, marking a significant leap in Apple silicon technology.

### Mistral OCR
- **Mistral OCR** is a cutting-edge Optical Character Recognition API that excels in understanding complex document elements, achieving **94.89% accuracy** in rigorous benchmark tests, outperforming competitors like Google Document AI and Azure OCR.

### 2024 ACM A.M. Turing Award
- **Andrew Barto and Richard Sutton** are awarded the **2024 ACM A.M. Turing Award** for their foundational work in **reinforcement learning**, which has significantly advanced AI technologies since the 1980s.

### QwQ-32B: Embracing the Power of Reinforcement Learning
- **QwQ-32B** leverages **Reinforcement Learning (RL)** to enhance reasoning capabilities, achieving performance comparable to models with significantly more parameters, such as **DeepSeek-R1** with 671 billion parameters, showcasing the effectiveness of RL in scaling model intelligence.

### Cognitive Behaviors That Enable Self-Improving Reasoners
- **Test-time inference** enhances language models' ability to tackle complex challenges, revealing that intrinsic properties like **verification, backtracking, subgoal setting, and backward chaining** are crucial for effective self-improvement.

### AMD Announces "Instella" Open-Source 3B Language Models
- **AMD's Instella** is a **fully open-source** language model with **3 billion parameters**, trained on **Instinct MI300X GPUs**, delivering performance comparable to leading models like Llama 3.2 and Gemma-2.

### CompressARC Performance
- **CompressARC** achieves **34.75% accuracy** on the training set and **20% on evaluation**, utilizing a novel approach that trains a neural network from scratch during inference without pretraining or datasets.

### Show HN: Beating Pokemon Red with RL and <10M Parameters
- **Reinforcement Learning (RL) has enabled an agent to beat Pokémon Red with a policy of less than 10 million parameters**, showcasing a significant reduction in complexity compared to previous models. This achievement highlights the potential of RL in solving complex gaming challenges, particularly in JRPGs, which require intricate decision-making and reasoning.

### The Tiny Star Explosions Powering Moore's Law
- **EUV lithography**, powered by **tin plasma explosions**, draws parallels to supernova physics, enabling advancements in semiconductor technology essential for modern electronics.

### Simple Explanation of LLMs
- **In 2023, ChatGPT achieved 100 million users faster than any app in the Web 2.0 era**, highlighting the rapid adoption of Large Language Models (LLMs) and the emergence of numerous competitors like Anthropic and HuggingFace, which now hosts **1.4 million models**.

### Using GRPO to Beat o1, o3-mini and R1 at "Temporal Clue"
- **GRPO** outperformed models like **R1**, **o1**, and **o3-mini** on the **Temporal Clue** reasoning task, achieving results comparable to **Sonnet 3.7** while being over **100x cheaper** at inference time.

### SepLLM: Accelerate LLMs by Compressing One Segment into One Separator
- **SepLLM** introduces a novel framework that significantly **accelerates inference** by compressing segments between special tokens, reducing computational demands without losing critical information.

### Training a Rust 1.5B Coder LM with Reinforcement Learning (GRPO)
- The **1.5B Coder LM** achieved a notable improvement, increasing build pass rates from **~60% to ~80%** and unit test success from **22% to 37%** after training with **GRPO** on **15k examples**.

### Beyond Relevance: Optimizing for Multiple Objectives in Search and Recommendations
- **Optimizing for multiple objectives** in search and recommendations enhances user experience by addressing diverse needs, moving beyond mere relevance to create personalized interactions.

### Cognitive Behaviors That Enable Language Model Self-Improvement: Analyzing Verification, Backtracking, Subgoals, and Backward Chaining
- **Four cognitive behaviors**— **double-checking**, **seeking background knowledge**, **step-back reasoning**, and **heuristic relaxation**—enable language models to enhance their reasoning capabilities autonomously, yielding significant performance gains across various reasoning tasks.
