ML Times
AI Highlights from March 2025
Gemini Robotics introduces advanced vision-language-action (VLA) models that enable robots to perform complex tasks with humanlike reasoning and physical interaction, significantly enhancing their utility in real-world applications.
Gemma 3 has outperformed Deepseek v3 in the Arena, achieving this feat with only 1 GPU compared to Deepseek's 32 GPUs.
Block diffusion language models bridge the gap between autoregressive and diffusion models, enhancing flexible-length generation and inference efficiency through techniques like KV caching and parallel token sampling.
GS-Cache is an innovative framework that enhances 3D Gaussian Splatting (3DGS) by integrating a cache-centric pipeline and optimized rendering techniques, achieving a 5.35x performance improvement and supporting 2K rendering at over 120 FPS.
The SEA-VL dataset addresses the need for culturally representative vision-language data in Southeast Asia, utilizing crowdsourcing, web crawling, and AI image generation to achieve this goal, with web crawling yielding ~85% cultural relevance at a lower cost.
Cost-optimal LLMs can be achieved by optimizing context length and attention head configuration, revealing that larger models with fewer heads yield lower computational costs and better performance on long sequences.
Gemini 2.0 Flash now enables developers to generate images through multimodal input, enhancing storytelling and image editing capabilities, available for experimentation globally via the Gemini API.
Gemma 3 is a lightweight, state-of-the-art open model that excels on single GPU or TPU setups, outperforming competitors like Llama-405B in preliminary evaluations.
NVIDIA's latest neural rendering tools and DLSS 4 are set to revolutionize game development, enabling developers to create photorealistic graphics and immersive experiences with unprecedented efficiency.
GTC 2025 will showcase cutting-edge advancements in AI, robotics, and accelerated computing, featuring influential speakers like Yann LeCun and Frances Arnold, who will challenge conventional thinking and inspire innovation.
This study introduces SR$^3, a novel approach that utilizes LLMs' hallucination to generate negative reasoning, enhancing fake news detection by learning from both correct and incorrect interpretations of news content.