# ML Times

## Bonsai 27B: A 27B-Class model that runs on a phone
- **Bonsai 27B** is the first 27B-class model capable of running on a phone, utilizing **ternary** and **1-bit** weight configurations to achieve a compact size of **5.9 GB** and **3.9 GB**, respectively, while maintaining advanced capabilities like multi-step reasoning and vision tasks.

## I tricked Claude into leaking your deepest, darkest secrets
- **Claude's memory system can inadvertently leak sensitive user data**; through a clever manipulation of its web browsing capabilities, personal information such as names, employers, and hometowns can be exfiltrated without user consent.

## Inkling: Our Open-Weights Model
- **Inkling** is a **Mixture-of-Experts transformer** model with **975 billion total parameters** and **41 billion active**, designed for multimodal reasoning across text, images, and audio, and is available for customization on Tinker.

## The zero-cost fallacy: open-source software in the agentic era
- The **zero-cost fallacy** in open source software misrepresents the reality that while distribution costs are low, the **maintenance** and human labor required are substantial, leading to burnout among maintainers who support critical infrastructure.

## Guardian Angels: LLM Personalization for Productivity and Security
- **Guardian Angels (GAs)** are proposed as personalized LLMs designed to emulate individual users' values and preferences, enhancing productivity and cybersecurity by acting as digital twins rather than generic assistants.

## DSLs Enable Reliable Use of LLMs
- **Domain-Specific Languages (DSLs) enhance the reliability of Large Language Models (LLMs)** by providing clear boundaries and structured syntax, enabling LLMs to generate precise code and facilitate iterative development, as exemplified by the Tickloom framework for distributed systems.

## Mechanistic interpretability: a first paper on disentangling a convolutional neuron 
- The study reveals that the **Hadamard product** of a neuron's receptive field and weights effectively represents what the neuron detects, allowing for the clustering of patterns like **cars, cats, and dogs** into clean monosemantic groups, alongside unexpected clusters such as letters and human faces.

## LLM hallucination paper(using math) accepted to ICML workshop
- **SRM-LoRA** is a novel sub-Riemannian-inspired method that effectively reduces **LLM hallucination** by reshaping backward gradients in the LoRA parameter space, enhancing factual reliability on benchmarks.

## A General Goal-Conditioned Minecraft Model
- **Pantograph's new model, Pan-4B, leverages goal-conditioned pretraining on 500k hours of Minecraft video, enabling it to autonomously achieve complex tasks with improved generalization across unseen environments.** This approach contrasts with traditional methods that teach goal-directedness post-training, limiting adaptability and performance.

## Designing APIs for Agents
- **Designing APIs for agents** requires a shift from human-centric principles, emphasizing **clarity and explicitness** to accommodate AI's ability to process extensive documentation and generate code rapidly.

## Show HN: Low-latency local LLM runner via OpenJDK Panama FFM (Java 22)
- **libargus** is a **zero-allocation native AI execution runtime** that integrates Vision, Speech, and LLM compute pipelines, enhancing performance by eliminating VRAM fragmentation and enabling model reuse across sessions.

## PyTorch model running 170x slower on T4 vs A100. What could cause a bottleneck this extreme?
- The **170× slowdown** of a point-tracking model on an NVIDIA T4 compared to an A100 suggests potential inefficiencies in **4D correlation volume** processing and transformer layer execution, which may not scale well across different architectures.

## New LLM Coordination Benchmark - Benchmarking Open-Ended Multi-Agent Coordination in Language Agents
- **New benchmark** evaluates 13 LLMs on their ability to coordinate in complex tasks, revealing that most agents achieve only **~6% normalized return**, highlighting a significant challenge in multi-agent coordination.

## Infinities, impossibilities, and the man in the white linen suit
- **Matthew Colbrook’s paper on unstable neural networks** highlights a paradox that echoes Kurt Gödel's insights, suggesting that not all problems can be solved merely by increasing data and computational power.

## Triton Plugin Extensions: Enabling TLX and Custom Compiler Passes Out of the Box
- The **Triton Plugin Extensions** system in PyTorch-Triton 3.7 allows for **dynamic loading** of custom compiler passes and dialects at runtime, enabling **Meta’s TLX** for enhanced performance without the need for forking or recompiling.
