ML Times
Main Articles
AI tooling has accelerated both malware creation and detection
- Exemplar: The rapid response to the LiteLLM supply chain attack on March 24, 2026, involved a full malware analysis and public disclosure within a single conversation.
TurboQuant: Redefining AI efficiency with extreme compression
- Introduces advanced quantization algorithms that achieve massive compression for large language models and vector search engines, enabling efficient memory usage without accuracy loss.
A.T.L.A.S achieves a remarkable 74.6% LiveCodeBench pass@1-v(k=3)
- Using a frozen 14B model on a single consumer GPU, significantly improving from previous versions due to innovative techniques like constraint-driven generation and self-verified iterative refinement.
Agentica SDK achieves a remarkable 36.08% score on ARC-AGI-3
- Surpassing competitors while completing 7 out of 25 games at a significantly lower cost.
Reco's AI-driven rewrite of JSONata, named gnata
- Achieved a 1,000x speedup on common expressions, resulting in a cost reduction of $500K/year.
HyperAgents: Self-referential self-improving agents
- Capable of optimizing any computable task, showcasing a novel approach to AI development.
Fast regex search: indexing text for agent tools
- Using trigram decomposition to create efficient indexes, significantly improving search speed in large codebases compared to traditional methods like
grep.
- Using trigram decomposition to create efficient indexes, significantly improving search speed in large codebases compared to traditional methods like
Chroma Context-1: A 20B parameter agentic search model
- Excels in multi-hop retrieval tasks with a self-editing context mechanism, managing information efficiently over extended searches.
6.4% of LoCoMo's answer key is incorrect
- The judge accepts up to 63% of intentionally wrong answers, highlighting significant flaws in the evaluation process.
Energy-Based Models (EBMs)
- Differ fundamentally from multi-layered perceptrons (MLPs) in their treatment of out-of-distribution (OOD) points, avoiding linear assumptions.
Qwen 3.5 27B achieved 1.1M tokens/second
- On 96 B200 GPUs using vLLM v0.18.0, significantly boosting throughput due to model size limitations.
Sup AI achieves an unprecedented 52.15% accuracy on the Humanity's Last Exam
- Outperforming its closest competitor by 7.41 points through innovative ensemble search and logprob scoring methods.
ARC-like data is crucial for model performance
- Well-performing models likely incorporate such data in their training sets, enhancing their reasoning capabilities.
CNN detection fails on MP3
- Findings reveal that CNNs trained on mel-spectrograms struggle with compressed audio, necessitating a dual-engine approach for effective detection of AI-generated music.
The Gumbel MCTS implementation in Python/Numba
- Achieves 2-15X faster performance than traditional PUCT, making it a significant advancement for MCTS applications.