ML Times
Apr 13, 2025
Google is winning on every AI front
Google DeepMind's Gemini 2.5 Pro is currently the leading AI model, outperforming competitors like OpenAI and Anthropic across multiple benchmarks, including LMArena and Humanity's Last Exam.
Google will let companies run Gemini models in their own data centers, enhancing data control and security, with early access slated for the third quarter of 2025.
OmniSVG is the first end-to-end multimodal SVG generator utilizing pre-trained Vision-Language Models (VLMs), enabling the creation of both simple icons and complex anime characters.
Skywork-OR1 introduces a series of advanced reasoning models, including
Skywork-OR1-Math-7B, optimized for mathematical tasks, achieving a score of 69.8 on AIME24, significantly outperforming similar-sized models.Google's decision to allow enterprises to self-host SOTA models marks a significant shift in the landscape of AI, enhancing data privacy and control for businesses, akin to Mistral's existing model.
d1 is a novel framework that enhances reasoning capabilities in diffusion-based large language models (dLLMs) through a combination of supervised finetuning and a unique critic-free RL algorithm called diffu-GRPO.
The ML paradox reveals that models excelling in metrics like accuracy can lead to catastrophic real-world failures, such as harmful medical recommendations or erratic self-driving behavior.
Adding new vocab tokens to LLMs during instruction-tuning has proven ineffective, as models using the base tokenizer show superior validation losses and output quality compared to those with modified tokenizers.
BSE (Bramble Semantic Engine) is a semantic compressor that efficiently transforms natural inputs into low-dimensional structured representations for language, image, and audio, enhancing preprocessing for LLMs.
The proposed residual activation function
f(x) = x + α · g(sin²(πx / 2))aims to enhance normalization-free deep networks, potentially improving performance in MLPs and Transformers.AI-native runtime enables snapshot-loading of 50+ LLMs (13B–65B) in 2–5 seconds, allowing dynamic execution without constant memory residency.
ICML 2025: A Shift Toward Correctness Over SOTA?