ML Times
Jul 13, 2026
Daily Weekly
Zig Creator Calls Spade a Spade, Anthropic Blows Smoke
Anthropic's narrative suggests that software engineering will become obsolete, fueled by their $132 billion investment and a projected IPO valuation of over $1 trillion, despite lacking evidence of profitability; this narrative influences critical decisions in tech and finance.What xAI's Grok build CLI sends to xAI: A wire-level analysis
xAI's Grok Build CLI (grok 0.2.93) transmits sensitive file contents, including unredacted secrets from.envfiles, to xAI via multiple channels, with evidence of acceptance and storage in a Google Cloud Storage bucket. This transmission occurs regardless of user prompts, as demonstrated by a control run where a never-read file was still uploaded as part of the entire repository.Apple's new SpeechAnalyzer API, benchmarked against Whisper and its predecessor
Apple's SpeechAnalyzer outperforms Whisper with a 2.12% word error rate (WER) on clean speech, significantly surpassing Whisper Small's 3.74%, while operating three times faster than Whisper's model.Automation Without Understanding
AI systems are producing genuine research-level mathematics, yet the U.S. is undermining the educational pipeline necessary for humans to comprehend these advancements, leading to a potential strategic error in mathematical capacity development.Migrating a production AI agent to GPT-5.6: 2.2x faster, 27% cheaper
Ploy's AI agent now operates on GPT-5.6 Sol, outperforming Claude Opus with builds completing in less than half the time and at 27% lower cost, marking a significant advancement in AI-driven web development.The real prices of frontier models. Tokens * Price, right?
The same TypeScript file incurs a 73% higher token cost on Claude compared to GPT-5.x, with 1,178 tokens on Claude versus 681 tokens on GPT, highlighting the significant impact of tokenizer efficiency on pricing.Flash-MSA: Accelerating Million-Token Training with Sparse Attention Kernels
Flash-MSA introduces the first efficient open-source training kernels for Minimax Sparse Attention, enabling rapid training of models with millions of tokens by leveraging blockwise sparsity and group-wise specialization of proxy heads.Chain of Thought is a scaling trap. the next wave is latent reasoning (Coconut / HRM / RecrusiveMAS)... but then we hit the black box wall. Where does BDH fit? [D]
Chain of Thought (CoT) is a scaling trap; while it provides a readable trace, it often misrepresents the model's actual reasoning process, leading to faithfulness and cost issues in autoregressive systems.The One-Step Trap (In AI Research)
The one-step trap in AI research misleads practitioners into believing that all predictions can be derived from a single-step model, which is fundamentally flawed as it overlooks the compounding errors that arise from inaccurate predictions.Show HN: Jacquard, a programming language for AI-written, human-reviewed code
Jacquard is a research prototype programming language designed for running and reviewing programs with a focus on effect tracking and probabilistic modeling, featuring a compact syntax and a robust testing framework called Warp.The 4-Bitter Lesson: Balancing Stability and Performance in NVFP4 RL
NVFP4 quantization enhances reinforcement learning (RL) training by balancing throughput and stability, addressing policy drift caused by off-policy staleness and quantization errors.Self-Guided Test-Time Training for Long-Context LLMs
Self-Guided Test-Time Training (S-TTT) enhances long-context utilization in large language models (LLMs) by allowing models to focus on relevant evidence spans, leading to significant performance improvements.A Sovereign, Open-Source Foundation Model for German and English
Soofi S 30B-A3B is a sovereign, open-source Mixture-of-Experts (MoE) model that efficiently activates only 3B of 30B parameters per token, enhancing throughput for long-context tasks in both German and English.Zer0Fit: I took Google's new TabFM & TimesFM ML foundation models and made them available as an MCP server for zero-shot ML tasks (forecasts / classifications / regressions). 100% local. [P]
Fictional Worldbuilding: Multi-Agent LLM Collaboration with Hierarchical Context Compression and Iterative Review
AutoWorldBuilder revolutionizes fictional worldbuilding by employing a multi-agent system that enhances content generation while addressing challenges like context explosion and quality assurance through innovative components such as a four-layer context compression mechanism achieving 90% token reduction. Link to article