Deep Research is now available on Gemini 2.5 Pro Experimental
Gemini 2.5 Pro Experimental enhances Deep Research, outperforming competitors in generating insightful reports, with user preference exceeding a 2-to-1 margin in tests.
Ironwood: The first Google TPU for the age of inference
Ironwood is Google’s seventh-generation TPU, engineered specifically for the age of inference, enabling advanced AI models to generate insights proactively rather than merely responding to queries.
No elephants: Breakthroughs in image generation
Multimodal image generation allows AI to create images directly, enhancing precision and coherence, as seen in the transition from traditional models that often misinterpret prompts, such as adding unwanted elements like elephants.
The Agent2Agent Protocol (A2A)
Agent2Agent (A2A) Protocol enables AI agents to communicate and collaborate across diverse platforms, enhancing productivity and reducing costs through interoperability among over 50 technology partners including Atlassian and Salesforce.
Visual Reasoning Is Coming Soon
OpenAI's GPT-4o model revolutionizes image manipulation by integrating full conversational context into image generation, allowing for more accurate and consistent results, such as dressing a specific cat in a detective hat.
[D] A regression head for llm works surprisingly well!
Introducing a regression head to a 33M VIT+decoder model significantly improved accuracy in visual grounding tasks, suggesting that this approach may enhance performance beyond traditional methods.
An LLM Query Understanding Service
LLMs can transform search query understanding by breaking down queries like "brown leather sofa" into structured intents, enabling rapid deployment of complex functionalities in days rather than months, all while running on local infrastructure without external API calls.
ProtoGS: Efficient and High-Quality Rendering with 3D Gaussian Prototypes
ProtoGS introduces a novel approach to 3D Gaussian Splatting by learning Gaussian prototypes, which significantly reduces the number of required Gaussian primitives while maintaining high visual quality.
[R] Beyond the Next Token: Towards Prompt-Robust Zero-Shot Classification via Efficient Multi-Token Prediction
P3 (Placeholding Parallel Prediction) enhances zero-shot text classification by predicting token probabilities across multiple positions, significantly improving model robustness and accuracy while reducing reliance on prompt engineering.
Re-Ranking in VPR: Outdated Trick or Still Useful? A study
Modern Visual Place Recognition (VPR) methods have advanced to a point where traditional image matching for re-ranking can actually degrade performance, suggesting a need for a paradigm shift in retrieval strategies.
[P] Reducing Transformer Training Time Without Sacrificing Accuracy — A Dynamic Architecture Update Approach
A novel method for optimizing transformer models enables dynamic architecture updates during training, significantly reducing training time while preserving accuracy.
Research Focus: Week of April 7, 2025
Microsoft Research unveils a global dataset of 375,197 wind turbines and 86,410 solar PV installations, aiding renewable energy planning by assessing ecological and cultural impacts through deep learning on 13 trillion pixels of satellite imagery.
From 128K to 4M: Efficient Training of Ultra-Long Context Large Language Models
This work presents an efficient training recipe that expands ultra-long context LLMs from 128K to 4M tokens, enhancing capabilities for applications like document understanding and inference-time scaling.
Lattice: Learning to Efficiently Compress the Memory
Lattice introduces a novel RNN mechanism that compresses memory using a fixed number of slots, achieving sub-quadratic complexity by leveraging the low-rank structure of K-V matrices.
FactGuard: Leveraging Multi-Agent Systems to Generate Answerable and Unanswerable Questions for Enhanced Long-Context LLM Extraction
FactGuard introduces a multi-agent system to autonomously generate both answerable and unanswerable questions, enhancing the training of long-context LLMs without the need for costly human annotation.