ML Times
Oct 16, 2025
ML Times Oct 16, 2025
Highlights:
Apple M5 chip
M5 achieves over 4x peak GPU compute performance for AI compared to M4, featuring a 10-core GPU with a Neural Accelerator in each core, enhancing both AI and graphics capabilities significantly.Claude Haiku 4.5
Claude Haiku 4.5 offers near-frontier coding performance at one-third the cost and twice the speed of its predecessor, Claude Sonnet 4, making it a game-changer for real-time AI applications.A Gemma model helped discover a new potential cancer therapy pathway
Google's C2S-Scale 27B model, with 27 billion parameters, has successfully identified a novel cancer therapy pathway by predicting the effects of silmitasertib in enhancing antigen presentation in immune-context-positive environments.New Coding Models and Integrations
Ollama introduces new coding models: The GLM-4.6 and Qwen3-Coder-480B are now available on Ollama’s cloud service, featuring seamless integration with familiar tools and enhanced performance for tool calling with Qwen3-Coder-30B.A kernel stack use-after-free: Exploiting Nvidia's GPU Linux drivers
Two critical vulnerabilities in NVIDIA's Linux GPU drivers, CVE-2025-23280 and CVE-2025-23300, allow local unprivileged processes to exploit kernel memory management, confirmed through a proof of concept that achieves kernel read and write primitives.State of AI Report 2025
The State of AI Report 2025 reveals that OpenAI maintains a slight edge in AI development, while China's DeepSeek and others are rapidly closing the gap in reasoning and coding tasks, marking a significant shift in global AI leadership.Recursive Language Models (RLMs)
Recursive Language Models (RLMs) enable language models to decompose and recursively interact with input contexts of unbounded length, significantly improving performance on long-context tasks while mitigating "context rot."SWE-Grep and SWE-Grep-Mini: RL for Fast Multi-Turn Context Retrieval
SWE-grep and SWE-grep-mini are newly trained models that achieve fast context retrieval in coding tasks, outperforming traditional models by an order of magnitude in speed while maintaining high accuracy.TaxCalcBench: Evaluating Frontier Models on the Tax Calculation Task
TaxCalcBench reveals that state-of-the-art models can accurately calculate less than one-third of federal income tax returns, highlighting significant limitations in current AI capabilities for tax filing.[P] Nanonets-OCR2: An Open-Source Image-to-Markdown Model with LaTeX, Tables, flowcharts, handwritten docs, checkboxes & More
Nanonets-OCR2 is a cutting-edge model suite that excels in converting images to markdown, featuring capabilities like LaTeX recognition, signature isolation, and multilingual support for diverse document types.[R]: Create a family of pre-trained LLMs of intermediate sizes from a single student-teacher pair
Boomerang distillation allows for the creation of a family of pre-trained LLMs of varying sizes by distilling a large teacher model into a smaller student and then reintegrating teacher layers, optimizing both performance and resource efficiency.Closer to production quality Python notebooks with
marimo check
marimo check is a linter designed to enhance the quality of notebooks, pipelines, and apps by providing actionable feedback for both humans and AI agents, ensuring adherence to best coding practices.Generalized Orders of Magnitude
Generalized Orders of Magnitude (GOOMs) extend traditional numerical methods, enabling stable computation over larger dynamic ranges than conventional floating-point approaches, crucial for fields like deep learning and finance.[R] Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Verbalized Sampling mitigates mode collapse in LLMs by prompting for probability distributions rather than single outputs, enhancing creative task diversity by 2.1x without sacrificing quality.PyTorch 2.9 Release Blog
PyTorch 2.9 introduces significant enhancements, including symmetric memory for multi-GPU programming and expanded support for AMD ROCm and Intel XPU, improving performance across diverse hardware platforms.