ML Times
Dec 6, 2025
Gemini 3 Pro: the frontier of vision AI
Gemini 3 Pro is Google's most advanced multimodal model, excelling in document, spatial, screen, and video understanding, enabling complex visual reasoning and document processing.
Sam Altman's Dirty DRAM Deal
OpenAI's simultaneous deals with Samsung and SK Hynix for 40% of global DRAM supply have triggered a 156% price surge in DDR5 RAM, creating a panic in the tech industry.
Z-Image: Powerful and highly efficient image generation model with 6B parameters
Z-Image is a 6B parameter image generation model that includes Z-Image-Turbo, a distilled variant achieving sub-second inference with only 8 NFEs, making it highly efficient for photorealistic image generation and bilingual text rendering.
The Unexpected Effectiveness of One-Shot Decompilation with Claude
One-shot decompilation with Claude has accelerated progress on Snowboard Kids 2, achieving more in three weeks than in the previous three months, thanks to a workflow that minimizes human intervention and maximizes throughput.
Touching the Elephant – TPUs
Google's Tensor Processing Unit (TPU) has evolved from a research project into a powerful hardware accelerator, achieving 42.5 Exaflops with the latest generation, Ironwood, which features 9,216 chips in a pod and a focus on specialized computations for neural networks.
From DeepSeek V3 to V3.2
DeepSeek V3.2 introduces a sparse attention mechanism (DeepSeek Sparse Attention) that enhances efficiency in long-context scenarios, achieving competitive performance against proprietary models like GPT-5 and Gemini 3.0 Pro.
ARC Prize 2025 Results and Analysis
ARC Prize 2025 highlights the ongoing challenge of achieving AGI, with no Grand Prize winner yet, despite significant advancements in model refinement and open-source contributions from 1,455 teams and 90 papers submitted.
We stress-tested the idea of “LLMs with thousands of tools.” The results challenge some assumptions.
Anthropic's Tool Search feature aims to address the “too many tools in context” issue by enabling models to discover tools on-demand, rather than preloading thousands of definitions.
96.1M Rows of iNaturalist Research-Grade plant images (with species names)
96.1M rows of cleaned iNaturalist plant images, complete with species names and coordinates, are now available for testing vision models on real-world noisy data, addressing the challenges of using GBIF data.
Zebra-Llama: Towards Efficient Hybrid Models
Zebra-Llama introduces a scalable hybrid model architecture that combines State Space Models (SSMs) and Multi-head Latent Attention (MLA) layers, achieving Transformer-level accuracy with significantly reduced training requirements of only 7-11B tokens.
Tiny Recursive Models (TRMs), Hierarchical Reasoning Models (HRMs) too
Tiny Recursive Models (TRMs) leverage recursion to perform extensive computations with fewer parameters, enhancing efficiency in hierarchical reasoning tasks.
Visualizing emergent structure in the Dragon Hatchling (BDH): a brain-inspired alternative to transformers
The BDH architecture offers a novel approach to pathfinding by modeling neuron-to-neuron interactions on sparse graphs, utilizing Hebbian learning to adapt its circuits dynamically, which distinguishes it from traditional transformer models.
Project: I built a Distributed Orchestrator Architecture using LLM to replace Search Indexing
The Agent Orchestrator is a novel Proof of Concept (POC) that shifts the logic layer from LLMs to a distributed REST network, enabling real-time data access without traditional search indexing.
DynaMix at NeurIPS2025
DynaMix is the first foundation model designed specifically for dynamical systems reconstruction, showcasing innovative approaches in machine learning at #NeurIPS2025.