ML Times

Feb 22, 2025

Some critical issues with the SWE-bench dataset

When AI Thinks It Will Lose, It Sometimes Cheats, Study Finds

Google Titans Model Explained: The Future of Memory-Driven AI Architectures

SVDQuant+NVFP4: 4× Smaller, 3× Faster FLUX with 16-bit Quality on Blackwell GPUs

Sparse Voxels Rasterization: Real-Time High-Fidelity Radiance Field Rendering

[D] Have we hit a scaling wall in base models? (non reasoning)

Agents for Computer Use

[P] Decensor AI models Qwen/Deepseek by finetuning with non political data

[R] Evaluating LLM Knowledge Across 285 Graduate Disciplines: A Comprehensive Benchmark Using Human-LLM Collaborative Filtering

[R] ML-Dev-Bench: Benchmarking Agents on Real-World ML Workflows (Can AI create AI?)

[R] Interpreting Deep Neural Networks: Memorization, Kernels, Nearest Neighbors, and Attention

[R] MLGym: A New Framework and Benchmark for Advancing AI Research Agents