ML Times
Jun 4, 2025
Daily
Deep learning gets the glory, deep fact checking gets ignored
Deep learning models, despite their glamour, can produce significant errors; a recent study revealed that a Transformer model made hundreds of incorrect predictions about enzyme functions, undermining its credibility.
Covert Web-to-App Tracking via Localhost on Android
- Meta and Yandex have developed a covert tracking method on Android that allows their apps to listen on local ports, enabling them to link web browsing data to user identities without consent. This method exploits the Android OS's permission model, allowing JavaScript from websites to communicate with native apps via localhost sockets, effectively bypassing privacy protections.
Cloud Run GPUs, now GA, makes running AI workloads easier for everyone
- Cloud Run GPUs are now generally available, enabling developers to run AI workloads with NVIDIA L4 GPUs without needing quota requests, thus simplifying access to GPU acceleration for serverless applications.
A deep dive into self-improving AI and the Darwin-Gödel Machine
- The Darwin-Gödel Machine (DGM) represents a significant advancement in self-improving AI, allowing systems to iteratively modify their own code based on empirical performance rather than formal proofs, thus mimicking biological evolution.
Vision Language Models are Biased
- Vision Language Models (VLMs) exhibit significant biases, scoring an average of 17.05% accuracy in counting tasks, revealing their inability to adapt to simple changes in visual stimuli, such as recognizing alterations in logos.
Doubling Down on Open Source
- Langfuse has open sourced all remaining Product Features under the MIT license, enhancing community collaboration and accelerating development cycles for LLM applications.
Human Brain Cells on Chip for Sale – First biocomputing platform hits the market
- Cortical Labs has launched the world's first biocomputer, utilizing human brain cells on a chip, with a price tag of $35,000 per unit, marking a significant advancement in biocomputing technology.
[R]Time Blindness: Why Video-Language Models Can't See What Humans Can?
- VLMs struggle with temporal patterns when spatial information is obscured, as demonstrated by the introduction of SpookyBench, where humans achieve over 98% accuracy while VLMs score 0%.
AGI is not multimodal
- AGI cannot be achieved through multimodal models as they fail to address the essential need for a physical understanding of the world, which is crucial for solving real-world problems like motion planning and social coordination.
Autonomous drone defeats human champions in racing first
- TU Delft's autonomous drone triumphed at the A2RL Drone Championship, marking a historic first by defeating human champions with speeds reaching 95.8 km/h on a challenging track, showcasing advanced AI capabilities in real-world conditions.
[R] Soft Thinking: Unlocking the Reasoning Potential of LLMs in Continuous Concept Space
- Soft Thinking introduces a novel method that enables human-like reasoning in LLMs by utilizing continuous concept space, allowing for smoother transitions and richer representations beyond discrete token limitations.
NVIDIA Blackwell Delivers Breakthrough Performance in Latest MLPerf Training Results
- NVIDIA Blackwell architecture achieved breakthrough performance in the latest MLPerf Training, outperforming previous generations by 2.2x on the Llama 3.1 405B pretraining benchmark, showcasing its capability to handle demanding AI workloads.
[N] Nvidia’s Blackwell Conquers Largest LLM Training Benchmark
- Nvidia's Blackwell GPUs have achieved unprecedented success in the latest MLPerf training results, outperforming competitors across all six benchmarks, while AMD's MI325X has shown competitive performance on the popular LLM fine-tuning benchmark.
🤗KV Cache from scratch in nanoVLM
- KV Caching in nanoVLM enhances autoregressive generation efficiency, achieving a 38% speedup by caching keys and values, thus avoiding redundant computations during token generation.
[R] GuidedQuant: Boost layer-wise PTQ methods using the end loss guidance (Qwen3, Gemma3, Llama3.3 / 2~4bit quantization) (ICML 2025)
- GuidedQuant enhances layer-wise PTQ methods by incorporating end loss guidance, leading to improved performance in quantization tasks across various models, including Qwen3 and Llama3.3.