ML Times
Nov 19, 2025
Gemini 3 is Google's most advanced AI model, enhancing reasoning and multimodal capabilities, enabling users to bring any idea to life across various Google products like the Gemini app and Vertex AI.
MAKER is the first system to solve a task with over one million LLM steps without errors, showcasing a breakthrough in long-range task execution.
Gemini 3 Pro enhances capabilities from Gemini 2.5, featuring 1 million input tokens and 64,000 output tokens, with improved performance in multimodal tasks including audio transcription and image processing.
MMaDA-Parallel introduces a parallel multimodal diffusion framework that enhances text-image generation by allowing continuous interaction between modalities, addressing performance degradation in complex tasks due to error propagation.
SAM 3 introduces advanced features in AI technology, enhancing user interaction and experience through improved algorithms and user interface design. For more details, visit the official article.
Mojo-V introduces a privacy-oriented RISC-V extension that enables secure, efficient secret computation by utilizing dedicated secret registers and third-party key encryption, achieving 5-7 orders of magnitude performance improvement over fully homomorphic encryption (FHE).
Segment Anything Model 3 (SAM 3) introduces a unified framework for detecting, segmenting, and tracking objects in images and videos using concept prompts, significantly enhancing Promptable Concept Segmentation (PCS) with a dataset of 4M unique labels.
Researchers from the University of Vienna uncovered a significant privacy vulnerability in WhatsApp's contact discovery mechanism, enabling the enumeration of 3.5 billion accounts across 245 countries, which has since been mitigated by Meta.
Microsoft, NVIDIA, and Anthropic have formed a strategic partnership to scale Anthropic's Claude AI model on Azure, with Anthropic committing to purchase $30 billion in Azure compute capacity and up to 1 gigawatt of additional capacity.
DeepClause is a neurosymbolic AI system that integrates symbolic reasoning with neural language models, enabling agents to perform complex logic and multi-step reasoning through its Prolog-based DSL, DML (DeepClause Meta Language).
Reproducible baselines for human action recognition achieve 87.05% accuracy on UCF-101 and 88.5% on Stanford40, providing a reliable starting point for researchers with pretrained models available on HuggingFace.
NVIDIA and Microsoft are enhancing AI capabilities by integrating next-generation technologies, including the NVIDIA Spectrum-X Ethernet switches and RTX PRO 6000 Blackwell GPUs, to power the new Microsoft Fairwater AI superfactory for large-scale AI training and inference.
Cornserve is an open-source platform designed to serve complex any-to-any multimodal AI models, such as Qwen 3 Omni, by utilizing a microservices architecture that enhances flexibility and scalability.