ML Times
May 11, 2025
Daily
Weekly
Absolute Zero: Reinforced Self-Play Reasoning with Zero Data
- Absolute Zero introduces a novel Reinforcement Learning with Verifiable Rewards (RLVR) paradigm that enables a model to autonomously generate tasks for its own learning, eliminating the need for external data and human supervision.
Venture Capital: In 2025, venture capital can't pretend everything is fine any more
- Venture capital is in crisis, with a heavy reliance on AI, particularly OpenAI, as the only viable investment, while other sectors languish in stagnation.
Wearable Technology: Engineers develop wearable heart attack detection technology
- Engineers at the University of Mississippi have developed a lightweight chip for wearable devices that detects heart attacks in real-time, potentially doubling the speed of traditional detection methods while maintaining 92.4% accuracy.
Adaptive Hashing
- Adaptive hashing enhances hash table performance by dynamically adjusting the hash function based on key distribution, leading to fewer collisions and improved cache efficiency.
Adventures in Imbalanced Learning and Class Weight
- Class weighting in imbalanced learning may not significantly enhance model performance, as empirical results suggest that optimal weights are only slightly above unweighted training, contradicting the common practice of using inverse proportion weighting.
The Evolution of RL for Fine-Tuning LLMs (from REINFORCE to VAPO)
- The evolution of reinforcement learning (RL) methods for fine-tuning large language models (LLMs) has progressed from classic techniques like PPO and REINFORCE to advanced methods such as GRPO, ReMax, and VAPO, incorporating innovations like reward shaping and token-level losses.
Simulating Bias with Bayesian Networks - Feedback wanted!
- The project simulates bias in machine learning using a Bayesian network to analyze recruitment, revealing that any observed bias stems from the models rather than the training data itself.
SDFs and the Fast Sweeping Algorithm in Jax
- The Fast Sweeping Method (FSM) efficiently solves the Eikonal equation in O(n) time, making it suitable for applications like signed distance functions (SDFs) in machine learning, where it represents 3D surfaces implicitly.
Exploring a New Hierarchical Swarm Optimization Model: Multiple Teams, Managers, and Meta-Memory for Faster and More Robust Convergence
- The new hierarchical swarm optimization model utilizes multiple teams managed by a "team manager" with meta-memory to enhance exploration efficiency and convergence speed in complex optimization tasks.
LeRobot Community Datasets: The “ImageNet” of Robotics — When and How?
- The LeRobot community datasets aim to create a diverse, community-driven benchmark for robotics, akin to ImageNet, by simplifying data collection and fostering collaboration among contributors.