ML Times
Mar 29, 2026
TurboQuant offers a revolutionary two-stage algorithm that compresses KV cache in AI models, achieving a 6x reduction in memory size without sacrificing accuracy, thus addressing the growing memory demands of large language models (LLMs).
Vibe coding, a term coined by Andrej Karpathy, involves deploying AI-generated code without thorough review, leading to significant failures, including a 6-hour outage at Amazon that resulted in 6.3 million lost orders.
Meta and Arm are collaborating to create the Arm AGI CPU, a groundbreaking data center processor tailored for AI workloads, enhancing performance and efficiency beyond traditional CPUs.
TurboQuant achieves 3.2× memory savings by compressing model weights with near-optimal distortion, offering a drop-in replacement for
nn.Linearthat maintains performance.The thermodynamics of computation is fundamentally flawed due to its reliance on idealizations that ignore thermal fluctuations, leading to erroneous conclusions about non-dissipative processes at molecular scales.
A benchmark was created to evaluate LLMs on 28 physics laws, generating adversarial questions that exploit common cognitive biases and unit confusions, ensuring a rigorous assessment through symbolic math rather than subjective judgment.
PentaNet introduces a pentanary quantization scheme {-2, -1, 0, +1, +2}, enhancing model capacity by 47% compared to ternary quantization, while maintaining zero-multiplier inference benefits through simple bit-shifting.
Litellm versions 1.82.7 and 1.82.8 were compromised, allowing a malicious .pth file to execute on every Python process start, scraping sensitive data like SSH keys and API credentials without requiring imports.
The first open-source implementation of Hebbian fast-weight write-back for the BDH architecture enables the model to rewrite its own decoder weights during inference, utilizing sparse activation codes as addresses, which was previously unimplemented publicly.
New hafnium oxide memristors could reduce AI energy consumption by up to 70% by mimicking the brain's efficient neural connections, enabling both storage and processing in a single location.
The implementation of TurboQuant in Python introduces a novel approach to quantization by utilizing random rotation to achieve optimal 1D quantization without the need for calibration data or dataset-specific tuning, making it applicable across various contexts.
LVFace, utilizing a Vision Transformer (ViT) backbone, reportedly outperforms ArcFace in facial recognition tasks, particularly in scenarios involving partially occluded faces like those wearing masks, as evidenced by its first-place finish in the MFR-Ongoing challenge.