ML Times
May 27, 2025
Daily
Most leading chatbots routinely exaggerate science findings
Up to 73% of leading chatbots produce inaccurate summaries of scientific findings, often exaggerating claims and misrepresenting the original texts, as revealed by a study from Utrecht University and Western University.Lossless video compression using Bloom filters
The new_bloom_filter_repo introduces a novel approach to lossless video compression using Rational Bloom Filters, which enhance traditional Bloom filters by allowing non-integer hash function counts for improved efficiency in data representation.TSMC bets on unorthodox optical tech
TSMC's innovative use of MicroLED-based interconnects aims to enhance energy efficiency in AI data centers, potentially revolutionizing data transmission methods.Mistral Agents API
The Mistral Agents API enhances AI capabilities by integrating persistent memory, code execution, and web search, enabling agents to perform complex tasks and maintain context across interactions.Revisiting the Algorithm That Changed Horse Race Betting
Bill Benter's horse betting algorithm, which generated over $1 billion in profits, utilizes a multinomial logit model to estimate winning probabilities, refined through decades of data and modern coding techniques.World's first petahertz transistor at ambient conditions
Researchers at the University of Arizona have developed the world's first petahertz-speed phototransistor, capable of operating in ambient conditions, which could enable computers to process data over 1,000 times faster than current technology.Outcome-Based Reinforcement Learning to Predict the Future
Outcome-based reinforcement learning (RL) can achieve frontier-scale accuracy in forecasting by adapting algorithms like Group-Relative Policy Optimisation (GRPO) and ReMax, effectively handling binary, delayed, and noisy rewards in real-world applications.Launch HN: Relace (YC W23) – Models for fast and reliable codegen
Relace offers a Fast Apply model that merges code snippets at 4300 tokens per second, significantly reducing merge errors compared to competitors like Sonnet and Llama, while also saving ~40% on Claude 4 output tokens.Show HN: My LLM CLI tool can run tools now, from Python code or plugins
LLM 0.26 introduces the ability for large language models to run tools directly in the terminal, allowing integration with models from OpenAI, Anthropic, and others through a Python function interface and CLI tool.[R] AutoThink: Adaptive reasoning technique that improves local LLM performance by 43% on GPQA-Diamond
AutoThink enhances local model reasoning by 43% on GPQA-Diamond through adaptive resource allocation, dynamically adjusting thinking time based on query complexity.[P] Evolving Text Compression Algorithms by Mutating Code with LLMs
Evolving text compression algorithms through mutations with LLMs achieved a compression ratio of 1.85, significantly improving from an initial ratio of 1.03 over 30 generations.Grammars of Formal Uncertainty
LLMs exhibit a significant accuracy variance in automated reasoning, with performance ranging from +34.8% on logical tasks to -44.5% on factual tasks, highlighting the need for careful evaluation of their outputs.[R] Panda: A pretrained forecast model for universal representation of chaotic dynamics
Panda is a pretrained model that forecasts chaotic dynamics, trained on a dataset of 20,000 chaotic systems, showcasing its ability to perform zero-shot predictions on real-world data without prior exposure.Demonstrating end-to-end scientific discovery with Robin: a multi-agent system
Robin, a multi-agent system, has successfully automated the entire scientific discovery process, leading to the identification of ripasudil as a novel treatment for dry age-related macular degeneration (dAMD), showcasing the potential of AI in drug discovery.Show HN: Maestro – A Framework to Orchestrate and Ground Competing AI Models
Maestro is a framework designed to orchestrate multiple large language models (LLMs) in parallel, allowing them to argue and mix outputs while preserving dissent, rather than selecting a single "best" response.