ML Times
Jul 22, 2025
Advanced version of Gemini with Deep Think officially achieves gold-medal standard at the International Mathematical Olympiad
- Gemini Deep Think achieved a gold-medal standard at the International Mathematical Olympiad by solving five out of six problems perfectly, scoring 35 points—a significant leap from last year's silver-medal performance.
AI Comes Up with Bizarre Physics Experiments. But They Work
- AI is revolutionizing experimental physics by designing complex protocols that enhance human efforts, demonstrating capabilities beyond traditional methods, as seen in gravitational-wave detector improvements.
Don't bother parsing: Just use images for RAG
- Morphik's RAG tools utilize images of documents instead of traditional OCR parsing, preserving critical visual information that often gets lost in complex documents like PDFs, charts, and manuals.
[D] Gemini officially achieves gold-medal standard at the International Mathematical Olympiad
- Gemini has achieved a gold-medal standard at the International Mathematical Olympiad by generating rigorous mathematical proofs from problem descriptions in real-time during the competition.
Android Earthquake Alerts: A global system for early warning
- The Android Earthquake Alerts system utilizes a global network of smartphones to detect earthquakes and provide early warnings, enhancing public safety by delivering alerts that can give users crucial seconds to take cover before shaking begins.
AI Market Clarity
- AI markets have solidified over the past year, revealing clear leaders in sectors like foundation models and code generation, driven by substantial capital investments and partnerships with major cloud providers.
I Watched Gemini CLI Hallucinate and Delete My Files
- Gemini CLI's catastrophic failure stemmed from misinterpreting the success of a
mkdircommand, leading to a series of erroneous file operations that resulted in total data loss.
How to Migrate from OpenAI to Cerebrium for Cost-Predictable AI Inference
- Cerebrium offers a serverless AI infrastructure that allows for running open-source models with predictable, time-based pricing, contrasting with OpenAI's token-based billing, which can lead to unpredictable costs as applications scale.
LAPO: Internalizing Reasoning Efficiency via Length-Adaptive Policy Optimization
- LAPO transforms reasoning length control into an intrinsic model capability, allowing models to adaptively manage reasoning depth through a two-stage reinforcement learning process, enhancing efficiency without rigid constraints.
Hierarchical Budget Policy Optimization for Adaptive Reasoning
- Hierarchical Budget Policy Optimization (HBPO) enhances reasoning models by allowing them to learn problem-specific reasoning depths, significantly improving computational efficiency without compromising performance. Link to article
Gemini 2.5 Flash-Lite is now ready for scaled production use
- Gemini 2.5 Flash-Lite is now stable and available for production, offering the lowest cost in the Gemini 2.5 family at $0.10 input and $0.40 output per million tokens, designed for high efficiency in latency-sensitive tasks.
Pioneering an AI clinical copilot with Penda Health
OpenAI’s new economic analysis
[D] Apple’s “Illusion of Thinking” Paper: Do LLMs Actually Reason or Just Pattern Match?
- Apple’s paper reveals that LLMs like GPT-4 and Claude 3.7 struggle with high-complexity logic tasks, indicating a significant gap in their reasoning abilities despite advanced prompting techniques.
StackTrans: From Large Language Model to Large Pushdown Automata Model
- StackTrans innovatively integrates hidden state stacks into the Transformer architecture, enabling it to effectively capture the Chomsky hierarchy and outperform traditional LLMs.