The Claude Code Source Leak: fake tools, frustration regexes, undercover mode
Anthropic's accidental source code leak reveals mechanisms like anti-distillation to inject fake tools, aimed at confusing competitors and protecting proprietary data, while also exposing a potential vulnerability in their deployment process.
Show HN: 1-Bit Bonsai, the First Commercially Viable 1-Bit LLMs
PrismML is pioneering ultra-dense AI models, such as the 1-bit Bonsai 8B, which requires only 1.15GB of memory, achieving 14× smaller footprint, 8× faster processing, and 5× less energy consumption compared to traditional models, while maintaining competitive performance on benchmarks.
Cohere Transcribe: Speech Recognition
Cohere Transcribe is a cutting-edge open-source automatic speech recognition (ASR) model, achieving an average word error rate (WER) of 5.42%, outperforming all competitors on the HuggingFace Open ASR Leaderboard.
From 300KB to 69KB per Token: How LLM Architectures Solve the KV Cache Problem
KV cache in AI models physically stores conversation data as bytes, drastically reducing computational redundancy by allowing new tokens to reference previously cached information, thus transforming memory management from quadratic to linear complexity.
🤗Falcon Perception
Falcon Perception enhances image-to-text capabilities, showcasing advanced OCR technology that improves accuracy and efficiency in data extraction from images.
[D] TurboQuant author replies on OpenReview
TurboQuant's novelty lies in its derivation of the exact distribution of rotated vector coordinates, which enables optimal coordinate-wise quantization, rather than merely exploiting existing distributional knowledge.
AI for American-produced cement and concrete
Meta's AI model, BOxCrete, enhances concrete mix design by utilizing Bayesian optimization to create sustainable, high-quality mixes exclusively from U.S. materials, addressing the significant reliance on imported cement.
TurboQuant KV Compression and SSD Expert Streaming for M5 Pro and IOS
SwiftLM is a native Swift inference server optimized for Apple Silicon, delivering OpenAI-compatible API performance without Python overhead, ensuring rapid model inference and efficient memory usage.
Inside the 'self-driving' lab revolution
AI-powered robotic tools like Eve are revolutionizing early-stage drug design by autonomously screening thousands of chemicals, significantly enhancing research efficiency and accuracy.
[D] Why I abandoned YOLO for safety critical plant/fungi identification. Closed-set classification is a silent failure mode
YOLO’s closed-set architecture fails to recognize out-of-distribution (OOD) inputs, leading to potentially dangerous misclassifications in safety-critical applications like plant and fungi identification, where accurate identification is crucial.
[R] Gram Newton-Schulz: A Fast, Hardware-Aware Newton-Schulz Algorithm for Muon
Gram Newton-Schulz optimizes the Newton-Schulz algorithm for Muon, achieving a 50% reduction in optimizer time for trillion-parameter models by iterating on a smaller Gram matrix instead of the original rectangular matrix, thus leveraging symmetric matrix multiplication efficiencies.
[P] EVōC: Embedding Vector Oriented Clustering
EVōC is a new library designed for clustering embedding vectors, addressing the challenges posed by high dimensionality that often hinder classical algorithms' performance.
🤗Granite 4.0 3B Vision: Compact Multimodal Intelligence for Enterprise Documents
Granite 4.0 3B Vision introduces a compact multimodal intelligence system designed to enhance the processing of enterprise documents, integrating image and text capabilities for improved efficiency.
Accelerating the next phase of AI
ADeLe: Predicting and explaining AI performance across tasks
ADeLe evaluates AI models by scoring 18 core abilities, enabling accurate predictions of performance on new tasks with ~88% accuracy, including models like GPT-4o and Llama-3.1.