# Apr 1, 2026

## Daily

### The Claude Code Source Leak: fake tools, frustration regexes, undercover mode
- **Anthropic's accidental source code leak** reveals mechanisms like **anti-distillation** to inject fake tools, aimed at confusing competitors and protecting proprietary data, while also exposing a potential vulnerability in their deployment process.

### Show HN: 1-Bit Bonsai, the First Commercially Viable 1-Bit LLMs
- **PrismML** is pioneering ultra-dense AI models, such as the **1-bit Bonsai 8B**, which requires only **1.15GB of memory**, achieving **14× smaller footprint**, **8× faster** processing, and **5× less energy consumption** compared to traditional models, while maintaining competitive performance on benchmarks.

### Cohere Transcribe: Speech Recognition
- **Cohere Transcribe** is a cutting-edge **open-source automatic speech recognition (ASR)** model, achieving an average **word error rate (WER) of 5.42%**, outperforming all competitors on the HuggingFace Open ASR Leaderboard.

### From 300KB to 69KB per Token: How LLM Architectures Solve the KV Cache Problem
- **KV cache** in AI models physically stores conversation data as bytes, drastically reducing computational redundancy by allowing new tokens to reference previously cached information, thus transforming memory management from quadratic to linear complexity.

### 🤗Falcon Perception
- **Falcon Perception** enhances image-to-text capabilities, showcasing advanced **OCR** technology that improves accuracy and efficiency in data extraction from images.

### [D] TurboQuant author replies on OpenReview
- **TurboQuant's novelty** lies in its derivation of the exact distribution of rotated vector coordinates, which enables **optimal coordinate-wise quantization**, rather than merely exploiting existing distributional knowledge.

### AI for American-produced cement and concrete
- **Meta's AI model, BOxCrete, enhances concrete mix design** by utilizing Bayesian optimization to create sustainable, high-quality mixes exclusively from U.S. materials, addressing the significant reliance on imported cement.

### TurboQuant KV Compression and SSD Expert Streaming for M5 Pro and IOS
- **SwiftLM** is a **native Swift inference server** optimized for Apple Silicon, delivering **OpenAI-compatible API** performance without Python overhead, ensuring rapid model inference and efficient memory usage.

### Inside the 'self-driving' lab revolution
- **AI-powered robotic tools** like Eve are revolutionizing early-stage drug design by autonomously screening thousands of chemicals, significantly enhancing research efficiency and accuracy.

### [D] Why I abandoned YOLO for safety critical plant/fungi identification. Closed-set classification is a silent failure mode
- **YOLO’s closed-set architecture fails to recognize out-of-distribution (OOD) inputs**, leading to potentially dangerous misclassifications in safety-critical applications like plant and fungi identification, where accurate identification is crucial.

### [R] Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon
- **Gram Newton-Schulz** optimizes the Newton-Schulz algorithm for Muon, achieving a **50% reduction in optimizer time** for trillion-parameter models by iterating on a smaller Gram matrix instead of the original rectangular matrix, thus leveraging symmetric matrix multiplication efficiencies.

### [P] EVōC: Embedding Vector Oriented Clustering
- **EVōC** is a new library designed for **clustering embedding vectors**, addressing the challenges posed by high dimensionality that often hinder classical algorithms' performance.

### 🤗Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents
- **Granite 4.0 3B Vision** introduces a **compact multimodal intelligence** system designed to enhance the processing of enterprise documents, integrating image and text capabilities for improved efficiency.

### Accelerating the next phase of AI

### ADeLe: Predicting and explaining AI performance across tasks
- **ADeLe** evaluates AI models by scoring **18 core abilities**, enabling accurate predictions of performance on new tasks with **~88% accuracy**, including models like GPT-4o and Llama-3.1.
