ML Times
Dec 22, 2025
2025 marked a transformative year
A year of vibes
2025 marked a transformative year for Armin Ronacher, as he transitioned from traditional programming to utilizing Claude Code and other AI tools, leading to a significant shift in his workflow and productivity.
Structured outputs create false confidence
- Structured outputs degrade response quality by forcing models to prioritize format over accuracy, leading to errors in data extraction and reasoning, as demonstrated with receipt parsing examples.
ONNX Runtime and CoreML May Silently Convert Your Model to FP16
- ONNX Runtime (ORT) with CoreMLExecutionProvider may convert models to FP16, leading to altered predictions; to maintain FP32 precision, specify the model format as MLProgram during inference session creation.
Universal Reasoning Model (53.8% pass 1 ARC1 and 16.0% ARC 2)
- The Universal Reasoning Model (URM) enhances universal transformers by integrating short convolution and truncated backpropagation, leading to significant performance improvements in reasoning tasks.
[R] EGGROLL: trained a model without backprop and found it generalized better
- EGGROLL demonstrates that optimizing NDCG directly using evolution strategies can outperform traditional contrastive loss methods, achieving a 22% improvement in validation scores despite a lower training score.
Why “negative vectors” can't delete data in FAISS – but weighted kernels can
- The Negative Vector Deletion Experiment demonstrates that inserting negated vectors into ANN indices does not achieve deletion, while operator-based data structures (OBDS) can accomplish true O(1) unlearning through destructive interference.
[R] Universal Reasoning Model
- The Universal Reasoning Model achieves 53.8% pass@1 on ARC-AGI 1 and 16.0% pass@1 on ARC-AGI 2, indicating a significant advancement in reasoning capabilities compared to previous models.
Continuously hardening ChatGPT Atlas against prompt injection
[P] A memory efficient TF-IDF project in Python to vectorize datasets larger than RAM
Learning What to Write: Write-Gated KV for Efficient Long-Context Inference
- Write-Gated KV introduces a novel approach to KV cache management, predicting token utility to enhance long-context inference efficiency, achieving 46-57% reduction in memory usage and 3.03-3.45× speedups on the Llama model.