# May 22, 2025

Daily

### Claude 4

- **Claude Opus 4** and **Sonnet 4** are the latest models from Anthropic, showcasing **unprecedented coding capabilities** and advanced reasoning, with Opus 4 leading in performance metrics like **72.5% on SWE-bench** and **43.2% on Terminal-bench**.

### Devstral

- **Devstral** is the leading open-source model for coding agents, achieving a score of **46.8% on SWE-Bench Verified**, surpassing previous models by over **6%** and outperforming larger models like Deepseek-V3-0324 (671B) and GPT-4.1-mini by more than **20%**.

### Gemini Diffusion

- **Gemini Diffusion** is Google's first LLM utilizing diffusion models, enhancing **speed** and **coherence** in text generation by refining noise rather than predicting text sequentially, achieving **857 tokens/second** in practical tests.

### For algorithms, a little memory outweighs a lot of time

- **Ryan Williams' groundbreaking proof** reveals that a small amount of memory can be as effective as extensive time in computational tasks, marking a significant advancement in complexity theory after 50 years of stagnation.

### LLM function calls don't scale; code orchestration is simpler, more effective

- **LLM function calls are inefficient**; using structured output schemas allows for **code orchestration**, which simplifies data processing and enhances scalability.

### An upgraded dev experience in Google AI Studio

- **Google AI Studio** now features **Gemini 2.5 Pro**, enhancing app development with native code generation and multimodal capabilities, allowing users to create applications from simple prompts and iterate through chat interactions.

### Google already out with a Text- Diffusion Model

- **Google's Gemini Diffusion** model introduces a novel approach to text generation, potentially enhancing reasoning capabilities beyond traditional transformer-based LLMs.

### µPC: Scaling Predictive Coding to 100 Layer Networks

- **$μ$PC enables the reliable training of 100+ layer predictive coding networks (PCNs)**, overcoming challenges that have historically hindered deep network performance, as demonstrated in recent studies (Pinchetti et al., 2024; Yang et al., 2023; Bordelon et al., 2023).

### Adventures in Symbolic Algebra with Model Context Protocol

- **MCP protocol** enables language models to **invoke external tools**, enhancing their capabilities beyond mere conversation, akin to a USB-C standard for AI integration.

### Loading Pydantic models from JSON without running out of memory

- **Pydantic's default JSON loading can consume up to 20× the size of the JSON file in memory**, making it impractical for large datasets, but using **`ijson`** for streaming parsing can significantly reduce this to just **1200MB**.

### Learn to Reason Efficiently with Adaptive Length-based Reward Shaping

- **LASER** introduces a **Length-bAsed StEp Reward shaping method** that optimizes reasoning efficiency in Large Reasoning Models (LRMs) by balancing performance and output length through adaptive rewards.

### Mini-satellite paves the way for quantum messaging anywhere on Earth

- A **Chinese team** has achieved a record in quantum communication by transmitting quantum-encrypted images over **12,900 kilometers**, demonstrating the potential for global quantum messaging using a **microsatellite**.

### Datadog releases SOTA time series foundation model and an observability benchmark

- **Datadog's Toto model** sets a new standard in time series forecasting, outperforming competitors on benchmarks like **BOOM**, **GIFT-Eval**, and **LSF** by leveraging proprietary observability data.

### The Annotated Kolmogorov-Arnold Network (Kan)

- The **Annotated Kolmogorov-Arnold Network (KAN)** offers a novel architecture that redefines activation functions by utilizing B-splines, enhancing interpretability and efficiency in deep learning models, while addressing limitations of traditional multi-layer perceptrons (MLPs).

### When to Continue Thinking: Adaptive Thinking Mode Switching for Efficient Reasoning

- **Adaptive Self-Recovery Reasoning (ASRR)** enhances efficiency in Large Reasoning Models (LRMs) by reducing unnecessary reasoning while maintaining performance, achieving up to **32.5%** reduction in reasoning budget with minimal accuracy loss.
