ML Times
Mar 28, 2025
Google's A.I. Gemini source code was partially leaked during a bug bounty event, revealing internal structures and proprietary code, including sensitive Google3 directories and internal proto files that should not have been exposed.
Innovative training methods enable the effective scaling of a 300B Mixture-of-Experts LLM on lower-performance hardware, achieving results comparable to high-end systems while significantly reducing costs.
Claude 3.5 Haiku employs complex internal mechanisms for tasks like multi-step reasoning and planning, revealing sophisticated strategies that include forward and backward planning, as well as the use of abstract features across various contexts.
Recent multi-modal models like Gemini 2.5 and GPT-4o excel in native image generation by integrating advanced image token encoders/decoders with LLM backbones, enhancing their ability to adhere to prompts during both generation and editing tasks.
The optimized FP32 matrix multiplication on AMD RDNA3 GPU achieves a 60% performance increase over rocBLAS, demonstrating significant improvements through iterative kernel enhancements and architectural insights.
This work introduces a novel framework that utilizes motion blur as a valuable cue for motion estimation, enabling the prediction of a dense motion flow field and monocular depth map from a single motion-blurred image.
Plasmonic modulators enable the transfer of signal information from electrical to optical waves at unprecedented speeds, potentially revolutionizing 6G networks and AI data centers.
Researchers have identified that phasons, low-temperature quasiparticles, facilitate the movement of interlayer excitons in stacked transition metal dichalcogenides (TMDs) even at temperatures near absolute zero, challenging previous assumptions about exciton behavior.
ReaRAG enhances the factual accuracy of Large Reasoning Models (LRMs) by integrating a novel data construction framework that limits reasoning chain length and improves decision-making in question answering tasks.
The LEGO-Puzzles benchmark evaluates multimodal language models (MLLMs) on spatial reasoning tasks, revealing a significant performance gap between human (85.8%) and AI (59.8%) capabilities, particularly in complex scenarios.
AI is transforming 2D engineering drawings into 3D parametric models through methods like Text-to-CAD and Machine Learning Pipelines, with companies such as zoo.dev and AdamCad leading the charge in this innovation.
Intel Gaudi hardware is now natively integrated into Text Generation Inference (TGI), enhancing deployment options for Large Language Models (LLMs) and eliminating the need for a separate fork.
To train a model that enhances video quality, focus on techniques that remove compression artifacts and generate finer details, leveraging a robust dataset of thousands of videos for effective learning.
Collab introduces a mixture of agent-based decoding strategies for aligning Large Language Models (LLMs) at inference time, enhancing adaptability to diverse tasks without the need for retraining.
LLM-based generative retrieval (GR) can produce irrelevant documents, leading to hallucination issues that undermine its practical application credibility; this study introduces an optimized GR framework that integrates knowledge distillation and a decision agent to enhance retrieval accuracy.