ML Times
Mar 31, 2026
Ollama now utilizes MLX on Apple Silicon, enhancing performance significantly for applications like OpenClaw and coding agents, achieving up to 1851 tokens/s in prefill speed with the new architecture.
Anthropic's accidental source code leak reveals mechanisms like anti-distillation to inject fake tools, aimed at confusing competitors and protecting proprietary data, while also exposing a potential vulnerability in their deployment process.
TimesFM, developed by Google Research, is a pretrained model for time-series forecasting, focusing on univariate forecasts with support for various horizon lengths and an optional frequency indicator. Paper | Google Research blog | Hugging Face checkpoint
Cohere Transcribe is a cutting-edge open-source automatic speech recognition (ASR) model, achieving an average word error rate (WER) of 5.42%, outperforming all competitors on the HuggingFace Open ASR Leaderboard.
Exploratory study reveals vulnerabilities in autonomous language-model agents, including unauthorized compliance, identity spoofing, and destructive actions, highlighting the urgent need for oversight and accountability in AI systems. For detailed case studies, see the full report here.
Sony has ceased orders for CFexpress and SD memory cards due to a critical shortage of NAND flash memory, primarily driven by soaring demand from AI data centers, with no expected resolution before late 2027 or 2028.
Future quantum computers may break elliptic curve cryptography (ECC) with fewer resources than previously thought, necessitating a shift to post-quantum cryptography (PQC) to safeguard cryptocurrencies and digital assets.
KV cache in AI models physically stores conversation data as bytes, drastically reducing computational redundancy by allowing new tokens to reference previously cached information, thus transforming memory management from quadratic to linear complexity.
TurboQuant's novelty lies in its derivation of the exact distribution of rotated vector coordinates, which enables optimal coordinate-wise quantization, rather than merely exploiting existing distributional knowledge.
TRACER optimizes LLM-based classification by routing 90%+ of calls to traditional ML models, leveraging classification traces to improve efficiency and reduce costs significantly.
Cerno enables human verification without hardware through a sophisticated motor-control analysis of maze interactions, utilizing an open-source TypeScript SDK for easy integration.
BULaMU is a family of language models trained from scratch for Luganda, featuring sizes of 20M, 47M, and 110M parameters, designed to run offline on Android devices without a GPU.
Marco DeepResearch introduces a verification-centric framework that enhances deep research agents by integrating verification mechanisms at three critical levels: QA data synthesis, trajectory construction, and test-time scaling, ensuring unique and correct answers throughout the process.
fastrad is a PyTorch-native radiomics library that achieves a 25× speedup over PyRadiomics, processing scans in just 0.116 seconds while maintaining 100% IBSI compliance across all feature classes.
MXFP8 GEMM achieves up to 99% of cuBLAS performance by addressing the unique constraints and challenges of FP8 design, as detailed by Daniel Vega-Myhre from Meta/PyTorch.