ML Times
Hierarchical Reasoning Model (HRM)
The Hierarchical Reasoning Model (HRM) introduces a novel recurrent architecture that enables efficient sequential reasoning with only 27 million parameters, achieving high performance on complex tasks without extensive data or pre-training. Link to article
PyroWave
PyroWave is a custom video codec designed for ultra-low latency game streaming, achieving encoding times as low as 0.13 ms on a RX 9070 XT, significantly outperforming traditional codecs like H.264 and HEVC.
CRISP Paper from Google DeepMind
The CRISP paper from Google DeepMind proposes integrating clustering during training, enhancing the model's ability to learn clusterable representations rather than relying on post-hoc methods, which are less effective.
X-pSRAM
X-pSRAM introduces a differential photonic SRAM bitcell that enables ultra-fast in-memory Boolean XOR computation, achieving operations at 10 GHz entirely in the optical domain, thus overcoming traditional bottlenecks in data movement.
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning
GEPA (Genetic-Pareto) leverages natural language reflection to optimize prompts, enabling LLMs to learn high-level rules more effectively than traditional reinforcement learning methods like GRPO, which often require extensive rollouts.
Misuse of ML for a Cortical Pain Biomarker
The critique in JAMA Neurology highlights methodological flaws in a previously published ML-based pain biomarker, specifically citing an incorrect validation set and an unrepresentative test set.
Extreme Single-Image Super-Resolution (SISR)
Extreme single-image super-resolution (SISR) techniques can achieve magnification factors of up to 100x, leveraging domain-specific texture synthesis tailored for materials.
Ambient Utils
ambient-utils is a Python package designed to train diffusion generative models using "bad data," focusing on denoising for specific diffusion times, enhancing model robustness.
AI-Failsafe-Overlay
The AI-Failsafe-Overlay introduces a logic-gated failsafe protocol aimed at addressing misalignment in recursive AI systems, featuring structural admission filters, audit-triggered lockdowns, and persistence-boundary constraints.
Step-3: Model-System Co-Design for Cost-Effective Decoding
Step-3 is a 321B-parameter VLM that employs a Multi-Matrix Factorization Attention (MFA) mechanism and Attention-FFN Disaggregation (AFD) to optimize decoding costs, achieving significant efficiency gains over existing models.
Smooth Reading
Smooth Reading introduces a chunk-wise inference method that enhances Recurrent LLMs' performance on long-context tasks by iteratively summarizing information, addressing their fixed-size memory limitations.
Utility-Based Passage Selector
Utility-based passage selection enhances retrieval-augmented generation (RAG) by focusing on the usefulness of passages rather than mere relevance, allowing for more accurate answers in complex queries.
PrismRAG
PrismRAG enhances retrieval-augmented generation (RAG) by integrating distractor-aware QA pairs and fostering reasoning skills, leading to improved model performance in complex contexts.
PyTorch on Kubernetes
Kubeflow Trainer is now integrated into the PyTorch ecosystem, providing a Kubernetes-native solution for scalable, distributed training of AI models, particularly for fine-tuning large language models (LLMs) with enhanced fault tolerance and resource management.