# May 28, 2025

## Daily

### Show HN: My LLM CLI tool can run tools now, from Python code or plugins
- **LLM 0.26 introduces the ability for large language models to run tools directly in the terminal**, allowing integration with models from OpenAI, Anthropic, and others through a Python function interface and CLI tool.

### Show HN: AutoThink – Boosts local LLM performance by 43% with adaptive reasoning
- **AutoThink** enhances local LLM performance by **43%** through adaptive reasoning, allocating computational resources based on **query complexity**—complex queries receive **70-90%** of tokens, while simple ones get **20-40%**.

### Look Ma, No Bubbles: Designing a Low-Latency Megakernel for Llama-1B
- **Low-latency megakernel** design for Llama-1B achieves **78% memory bandwidth utilization** on H100 GPUs, outperforming existing systems by over **1.5x** by merging multiple kernels into a single execution unit.

### [R] Bloat in machine learning shared libs is >70%
- The paper "The Hidden Bloat in Machine Learning Systems" reveals that **device code** is a major contributor to **bloat**, with reductions in code size of up to **75%** for device code and **72%** for host code using the tool **Negativa-ML**.

### Revisiting the algorithm that changed horse race betting (2023)
- **Bill Benter's horse betting algorithm**, which generated over **$1 billion** in profits, utilizes a **multinomial logit model** to estimate winning probabilities, refined through decades of data and modern coding techniques.

### Compiling a Neural Net to C for a 1,744× speedup
- **Compiling a neural network to C achieved a remarkable 1,744× speedup** in inference speed by utilizing logic gates instead of traditional activation functions, specifically for a 3×3 kernel function in Conway’s Game of Life.

### Developing first petahertz-speed phototransistor in ambient conditions
- Researchers at the University of Arizona have developed the **world's first petahertz-speed phototransistor**, capable of operating in ambient conditions, which could enable computers to process data over **1,000 times faster** than current technology.

### Launch HN: Relace (YC W23) – Models for fast and reliable codegen
- **Relace** offers a **Fast Apply model** that merges code snippets at **4300 tokens per second**, significantly reducing merge errors compared to competitors like Sonnet and Llama, while also saving **~40% on Claude 4 output tokens**.

### [R] New ICML25 paper: Train and fine-tune large models faster than Adam while using only a fraction of the memory, with guarantees!
- The new paper, **[Lean and Mean Adaptive Optimization via Subset-Norm and Subspace-Momentum with Convergence Guarantees](https://arxiv.org/abs/2411.07120)**, presents techniques that achieve **80% memory reduction** while maintaining performance comparable to Adam, using only **half the training tokens** for LLaMA 1B.

### Outcome-Based Reinforcement Learning to Predict the Future
- **Outcome-based reinforcement learning (RL) can achieve frontier-scale accuracy** in forecasting by adapting algorithms like **Group-Relative Policy Optimisation (GRPO)** and **ReMax**, effectively handling binary, delayed, and noisy rewards in real-world applications.

### [R] AutoThink: Adaptive reasoning technique that improves local LLM performance by 43% on GPQA-Diamond
- **AutoThink** enhances local model reasoning by **43%** on GPQA-Diamond through adaptive resource allocation, dynamically adjusting thinking time based on query complexity.

### [R] Grammars of Formal Uncertainty: When to Trust LLMs in Automated Reasoning Tasks
- **LLMs** exhibit a significant **accuracy variance** in automated reasoning, with performance ranging from **+34.8%** on logical tasks to **-44.5%** on factual tasks, highlighting the need for careful evaluation of their outputs.

### FlowTSE: Target Speaker Extraction with Flow Matching
- **FlowTSE** introduces a **novel** approach to **target speaker extraction (TSE)** using **conditional flow matching**, simplifying the process compared to existing generative methods that often involve complex pipelines.

### Direct Preference Optimization vs. RLHF
- **Direct Preference Optimization (DPO)** enhances language models by aligning them with human preferences through direct training on preference data, eliminating the need for complex reinforcement learning methods.

### Launch HN: MindFort (YC X25) – AI agents for continuous pentesting
- **MindFort** develops **autonomous AI agents** that continuously identify, validate, and patch security vulnerabilities in web applications, functioning as a **24/7 AI red team**.
