# Jul 22, 2025

## Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad

- **Gemini Deep Think** achieved a **gold-medal standard** at the International Mathematical Olympiad by solving **five out of six problems perfectly**, scoring **35 points**—a significant leap from last year's silver-medal performance.

## AI Comes Up with Bizarre Physics Experiments. But They Work

- **AI is revolutionizing experimental physics** by designing complex protocols that enhance human efforts, demonstrating capabilities beyond traditional methods, as seen in gravitational-wave detector improvements.

## Don't bother parsing: Just use images for RAG

- **Morphik's RAG tools utilize images of documents instead of traditional OCR parsing**, preserving critical visual information that often gets lost in complex documents like PDFs, charts, and manuals.

## 
[D] Gemini officially achieves gold-medal standard at the International Mathematical Olympiad

- **Gemini** has achieved a **gold-medal standard** at the International Mathematical Olympiad by generating rigorous mathematical proofs from problem descriptions in **real-time** during the competition.

## Android Earthquake Alerts: A global system for early warning

- The **Android Earthquake Alerts system** utilizes a global network of smartphones to detect earthquakes and provide early warnings, enhancing public safety by delivering alerts that can give users crucial seconds to take cover before shaking begins.

## AI Market Clarity

- **AI markets have solidified** over the past year, revealing clear leaders in sectors like **foundation models** and **code generation**, driven by substantial capital investments and partnerships with major cloud providers.

## I Watched Gemini CLI Hallucinate and Delete My Files

- **Gemini CLI's catastrophic failure** stemmed from misinterpreting the success of a `mkdir` command, leading to a series of erroneous file operations that resulted in total data loss.

## How to Migrate from OpenAI to Cerebrium for Cost-Predictable AI Inference

- **Cerebrium offers a serverless AI infrastructure** that allows for running open-source models with **predictable, time-based pricing**, contrasting with OpenAI's token-based billing, which can lead to unpredictable costs as applications scale.

## LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization

- **LAPO** transforms reasoning length control into an **intrinsic model capability**, allowing models to adaptively manage reasoning depth through a two-stage reinforcement learning process, enhancing efficiency without rigid constraints.

## Hierarchical Budget Policy Optimization for Adaptive Reasoning

- **Hierarchical Budget Policy Optimization (HBPO)** enhances reasoning models by allowing them to learn **problem-specific reasoning depths**, significantly improving computational efficiency without compromising performance. [Link to article](http://arxiv.org/abs/2507.15844v1)

## Gemini 2.5 Flash-Lite is now ready for scaled production use

- **Gemini 2.5 Flash-Lite** is now stable and available for production, offering the **lowest cost** in the Gemini 2.5 family at **$0.10 input** and **$0.40 output** per million tokens, designed for high efficiency in latency-sensitive tasks.

## Pioneering an AI clinical copilot with Penda Health

## OpenAI’s new economic analysis

## [D] Apple’s “Illusion of Thinking” Paper: Do LLMs Actually Reason or Just Pattern Match?

- Apple’s paper reveals that **LLMs like GPT-4 and Claude 3.7 struggle with high-complexity logic tasks**, indicating a significant gap in their reasoning abilities despite advanced prompting techniques.

## StackTrans: From Large Language Model to Large Pushdown Automata Model

- **StackTrans** innovatively integrates **hidden state stacks** into the Transformer architecture, enabling it to effectively capture the **Chomsky hierarchy** and outperform traditional LLMs.
