ML Times
Feb 7, 2026
Daily
The Waymo World Model
- The Waymo World Model introduces a generative model that enhances autonomous driving simulations by creating hyper-realistic environments, enabling the Waymo Driver to navigate complex scenarios before encountering them in real life.
Reinforcement Learning from Human Feedback
- Reinforcement Learning from Human Feedback (RLHF) is a pivotal method for enhancing machine learning systems, integrating insights from economics, philosophy, and optimal control to improve language models.
Monty: A minimal, secure Python interpreter written in Rust for use by AI
- Monty is a minimal, secure Python interpreter built in Rust, designed to execute LLM-generated code with startup times under 1μs, significantly reducing latency compared to traditional container-based sandboxes.
How we made geo joins 400× faster with H3 indexes
- Floe's innovative use of H3 indexes transforms geo joins, achieving a 400× speedup by converting complex spatial predicates into efficient set operations, allowing for rapid equi-joins on compact keys.
Software Factories and the Agentic Moment
- StrongDM's Software Factory utilizes non-interactive development, where agents autonomously write and validate code without human intervention, marking a significant shift in software engineering practices.
Show HN: Smooth CLI – Token-efficient browser for AI agents
- Smooth CLI revolutionizes AI agent interaction by providing a natural language interface, allowing agents to articulate goals rather than execute low-level commands, thus enhancing efficiency and focus.
Evaluating and mitigating the growing risk of LLM-discovered 0-days
- Claude Opus 4.6 demonstrates a significant leap in AI's ability to discover high-severity vulnerabilities in code, outperforming traditional fuzzing methods by reasoning like a human researcher rather than relying solely on random input testing.
FORTH? Really!?
- FORTH and associative languages may enhance transformer architectures by promoting a concatenative approach rather than a recursive one, suggesting a shift in how we structure problem-solving in AI.
First Proof
- The study presents ten unique math questions derived from the authors' research, aimed at evaluating the capabilities of current AI systems in solving complex mathematical problems. Link to article
StrongDM's AI team build serious software without even looking at the code
- StrongDM’s AI team has pioneered a Software Factory model where coding agents autonomously generate and validate code without human intervention, leveraging advanced AI capabilities to enhance software development efficiency.
[P] Wrote a VLM from scratch! (VIT-base + Q-Former + LORA finetuning)
- The author successfully finetuned a text-only language model into a vision language model (VLM) using a VIT-base encoder and a Q-Former model, achieving notable results in just four hours of training on a single V100 GPU for a minimal cost of 50 cents.
[R] Mixture-of-Models routing beats single LLMs on SWE-Bench via task specialization
- Mixture-of-Models architecture outperforms single LLMs on SWE-Bench by leveraging task-level specialization, allowing models to excel in specific subsets of tasks rather than relying on a single aggregate model.
Selection Rather Than Prediction
- Selection over prediction enhances coding efficiency by generating multiple candidate implementations, allowing for optimization rather than relying on a single agent's performance, which can be misleading due to high variance across tasks and languages.
Make Trust Irrelevant: A Gamer's Take on Agentic AI Safety
- Agentic AI safety fails because it focuses on making agents trustworthy rather than eliminating the need for trust, emphasizing that mechanics must govern actions, not intentions.
[P] How do you regression-test ML systems when correctness is fuzzy? (OSS tool)
- Regression testing for ML systems is challenging due to the absence of a single correct answer, leading to reliance on subjective evaluations of system behavior rather than traditional methods.