ML Times
Feb 18, 2026
Claude Sonnet 4.6
- Claude Sonnet 4.6 is a significant upgrade, featuring enhanced skills in coding, long-context reasoning, and a 1M token context window, making it the most capable model yet from Anthropic.
Advice, not control: the role of Remote Assistance in Waymo's operations
- Waymo's 6th-generation Driver marks a pivotal advancement in fully autonomous operations, enabling expansion into diverse environments, including extreme winter conditions, while reducing costs and upholding safety standards.
Testing INT8 Model on Snapdragon Chipsets
- [D] We tested the same INT8 model on 5 Snapdragon chipsets. Accuracy ranged from 93% to 71%. Same weights, same ONNX file.
- Accuracy varied significantly across five Snapdragon chipsets, with results ranging from 93% to 71%, highlighting the impact of hardware on model performance despite identical configurations.
The Future of AI Software Development
- The Thoughtworks Future of Software Development Retreat revealed that AI-assisted work is straining traditional software development practices, necessitating new frameworks and roles, as detailed in a 17-page summary.
OpenAI, the US government, and Persona built an identity surveillance machine
- OpenAI, Persona, and the US government have developed a surveillance system that utilizes facial recognition and identity verification to monitor users, filing reports on them to federal agencies. This system operates by scoring selfies against watchlists and continuously re-screening users, raising significant privacy concerns.
SkyRL brings Tinker to your GPUs (2025)
- Notion is a versatile productivity tool that integrates note-taking, task management, and collaboration features, enhancing workflow efficiency for individuals and teams alike.
Learning State-Tracking from Code Using Linear RNNs
- Linear RNNs demonstrate superior performance in state-tracking tasks, particularly in permutation composition, by converting these tasks into code via REPL traces, which reveal states through prints and variable transformations.
Data Lineage in ML Pipelines
- Many ML teams lack systematic data lineage tracking, often relying on manual methods or none at all, which complicates reproducibility and compliance with regulations like the EU AI Act (Article 10) that mandates data documentation for high-risk AI systems.
Run LLMs locally in Flutter with <200ms latency
- Edge-Veda is a managed on-device AI runtime for Flutter, enabling sustainable operation of text, vision, and speech models without cloud dependencies, ensuring privacy and stability under real-world constraints.
How ZeRO-1 could be faster than ZeRO-2?
- ZeRO-1 may outperform ZeRO-2 due to its efficient gradient storage strategy, which avoids unnecessary duplication across nodes while maintaining the same communication requirements.
STAPO: Stabilizing Reinforcement Learning for LLMs by Silencing Rare Spurious Tokens
- STAPO introduces a novel approach to stabilizing reinforcement learning in large language models by addressing the impact of spurious tokens, which constitute only 0.01% of the data yet cause significant instability in training.
Avey-B
- Avey emerges as a powerful autoregressive, attention-free alternative to traditional bidirectional encoders, offering a promising encoder-only adaptation that enhances efficiency in NLP tasks.
India Fuels Its AI Mission With NVIDIA
- India is investing over $1 billion in its AI ecosystem, partnering with NVIDIA to enhance compute capacity and develop sovereign AI datasets and models, crucial for its ambitious AI transformation goals.
NVIDIA Nemotron 2 Nano 9B Japanese
- NVIDIA's Nemotron 2 Nano 9B Japanese achieves state-of-the-art performance in the sub-10B parameter category, enhancing Japanese language understanding and agent capabilities for enterprise AI applications.
IBM and UC Berkeley Diagnose Why Enterprise Agents Fail Using IT-Bench and MAST
- IBM and UC Berkeley's study reveals that frontier models like Gemini-3-Flash exhibit isolated failure modes (2.6 per trace), while larger models like GPT-OSS-120B suffer from cascading failures (5.3 per trace), indicating a critical difference in reliability.