ML Times
Oct 2, 2025
Daily
Building the heap: racking 30 petabytes of hard drives for pretraining
- 30 petabytes of data storage were built for under $500,000 in San Francisco, enabling the pretraining of models on video data, which requires 500 times more storage than text-based models like LLaMa-405B.
We Bought the Whole GPU, So We're Damn Well Going to Use the Whole GPU
- New megakernel for tensor-parallel inference with Llama-70B on H100s achieves >22% higher throughput than SGLang by optimizing resource usage across compute, memory, and communication operations.
Unix philosophy and filesystem access makes Claude Code amazing
- Claude Code merges a terminal-based Unix command interface with filesystem access, enabling LLMs to possess persistent memory and seamless tool chaining, thus evolving into a robust agentic operating system for coding and note-taking.
OpenTSLM: Language models that understand time series
- OpenTSLM introduces Time-Series Language Models (TSLMs), which enable AI to reason about temporal data alongside traditional modalities, enhancing forecasting and explanation capabilities.
The RAG Obituary: Killed by agents, buried by context windows
- RAG's decline is imminent as context windows expand and agent-based architectures evolve, rendering traditional retrieval-augmented generation systems obsolete in handling complex documents.
Announcing Tinker
- Tinker is a flexible API designed for fine-tuning language models, enabling researchers to customize algorithms and data while simplifying distributed training complexities.
DARPA project for automated translation from C to Rust (2024)
- DARPA's TRACTOR program aims to automate the translation of legacy C code to the safer Rust language, addressing the pervasive issue of memory safety vulnerabilities that have plagued software for decades.
The Atlantic Quantum team is joining Google
- Google Quantum AI is accelerating its quantum computing capabilities by acquiring Atlantic Quantum, a startup known for its innovative modular chip stack that integrates qubits with superconducting control electronics.
Reverse-engineering Flash Attention 4
- Flash Attention 4 (FA4) enhances Transformer performance, achieving up to 22% faster execution than NVIDIA's cuDNN library through advanced CUDA kernels designed for attention layers.
Looking for papers, preprints, datasets, or reports where an LLM is trained to only know what humans knew before a major scientific breakthrough.
- Training an LLM on historical physics texts up to 1904 could yield novel theoretical frameworks, challenging the model to generate ideas without modern scientific knowledge or experimental validation.
Looking for a paper I saw once about training task solving models that output human readable explanations
- The paper discusses a two-model approach where the first model generates human-readable explanations from task encodings, while the second model utilizes these explanations to solve tasks, potentially enhancing interpretability in AI systems.
SOTA OCR on-device with Core ML and dots.ocr
- dots.ocr, a 3B parameter OCR model, outperforms Gemini 2.5 Pro in OmniDocBench, enabling competitive on-device OCR capabilities without network reliance or API key management.