ML Times
May 4, 2026
AI systems demonstrated superior accuracy in emergency triage
Diagnosing correctly in 67% of cases compared to 50-55% for human doctors, particularly excelling in high-pressure situations with limited information.
DeepClaude utilizes DeepSeek V4 Pro as a cost-effective alternative to Claude Code, achieving a 96.4% score on LiveCodeBench while reducing costs from $200/month to approximately $20/month for light usage.
SSMs exhibit a 3.26x compression disadvantage compared to attention QKV, severely impacting their performance in parameter-constrained environments like the Parameter Golf competition.
**NVENC silicon on modern GPUs can be leveraged to achieve approximately 180 GB/s effective bandwidth for cross-GPU communication, effectively replacing NVLink on consumer cards like the RTX 4090 and 5090.
The demand for privacy-preserving AI/ML has surged alongside the rise of LLMs, driven by concerns over user de-anonymization as highlighted in this paper.
Running LLMs locally or on private cloud is wasteful, as large batching during inference significantly enhances efficiency through optimized memory and compute scaling.
QLoRA fine-tuning of Qwen2.5-1.5B achieved an impressive accuracy of 84.9% in classifying English texts across six CEFR levels (A1–C2), utilizing only ~0.28% of model parameters for training.
AdaMeZO is a novel zeroth-order optimizer that achieves Adam-style performance without the memory overhead of maintaining moment estimates, thus enabling efficient fine-tuning of LLMs.
LLMs face unique challenges in information retrieval due to their limited attention budgets and susceptibility to noise, which can lead to hallucinations and reasoning failures, necessitating a focus on denoising to enhance evidence density and verifiability.
Budget-aware routing for long clinical text addresses the challenge of token cost in large language models by selecting document units under strict token budgets, optimizing for relevance, coverage, and diversity through a proposed method called RCD.