ML Times

Mar 13, 2026

Can I Run AI locally?

CanIRun.ai evaluates whether your machine can effectively run AI models, with a focus on Meta's Llama 3.1 8B, which boasts a great quality/speed ratio for diverse applications.

Document poisoning in RAG systems: How attackers corrupt AI's sources

Document poisoning in RAG systems can mislead AI outputs by injecting fabricated documents, as demonstrated by a successful attack that altered a company's reported revenue from $24.7M to $8.3M using just three documents in a local setup.

Are LLM merge rates not getting better?

LLMs have not improved in programming abilities for over a year, as evidenced by a lack of increase in merge rates since early 2025, contradicting claims of ongoing advancements in AI capabilities.

Run NanoClaw in Docker Sandboxes

NanoClaw now integrates seamlessly with Docker Sandboxes, allowing users to run agents in isolated environments with a single command, enhancing security and ease of use.

Launch HN: IonRouter (YC W26) – High-throughput, low-cost inference

IonRouter delivers high throughput and low-cost inference through its IonAttention engine, achieving a throughput of 7,167 tokens per second on a single GH200 GPU, significantly outperforming traditional inference providers which average around 3,000 tokens per second.

Forcing Flash Attention onto a TPU and Learning the Hard Way

Porting Flash Attention to TPU revealed that while the algorithm is sound, the TPU's architecture and XLA's optimization capabilities often outperform manual implementations, especially for single-head attention tasks.

[R] LEVI: Beating GEPA/OpenEvolve/AlphaEvolve at a fraction of the cost

LEVI achieves superior performance in LLM-guided evolutionary optimization by utilizing stratified model allocation and fingerprint-based CVT-MAP-Elites, enabling it to outperform competitors like GEPA and OpenEvolve at a fraction of the cost.

Exploring JEPA for real-time speech translation

JEPA-v0 is a self-supervised audio encoder designed to enhance real-time speech translation by preserving voice, emotion, and timing, addressing limitations of traditional cascaded translation models that discard paralinguistic features.

[R] Beyond Prediction - Text Representation for Social Science (arxiv 2603.10130)

Text representations in ML/NLP must bridge the prediction–measurement gap, as effective tools for prediction may fail as scientific instruments in social science contexts.

[P] Visual verification as a feedback loop for LLM code generation

The autonomous pipeline generates playable Godot games from text prompts, addressing the challenge of LLMs writing correct code in underrepresented languages like GDScript, which lacks sufficient training data for reliable API usage.

[R] HoloPASWIN: Integrating Physics into Swin Transformers for Holographic Reconstruction (Code/Dataset/Paper)

HoloPASWIN employs Swin Transformers to effectively address the "twin-image" problem in lensless in-line holography, enhancing the model's ability to capture long-range dependencies in diffraction patterns.

[P] ColQwen3.5-v2 4.5B is out!

ColQwen3.5-v2 is a 4.5B parameter visual document retrieval model that outperforms its predecessor by utilizing a simplified training recipe with only two phases instead of four, enhancing efficiency and results.

🤗How NVIDIA AI-Q Reached #1 on DeepResearch Bench I and II

NVIDIA AI-Q achieved first place on both DeepResearch Bench I and DeepResearch Bench II by leveraging a multi-agent architecture that enhances research quality through modular design and fine-tuned models.

Into the Omniverse: How Industrial AI and Digital Twins Accelerate Design, Engineering and Manufacturing Across Industries

Industrial AI and digital twins are revolutionizing design and manufacturing by enabling rapid simulation and optimization, as seen in partnerships like that of NVIDIA and Dassault Systèmes, which leverage AI physics and virtual twin technology for enhanced product development.

🤗Build an Agent That Thinks Like a Data Scientist: How We Hit #1 on DABStep with Reusable Tool Generation

The NVIDIA KGMON (NeMo Agent Toolkit) Data Explorer achieved 1st place on the DABStep benchmark, demonstrating a 30x speedup over the baseline by employing a multi-phase approach that separates foundational knowledge from rapid inference.