# Mar 29, 2026

- **TurboQuant** offers a revolutionary two-stage algorithm that compresses KV cache in AI models, achieving a **6x reduction in memory size** without sacrificing accuracy, thus addressing the growing memory demands of large language models (LLMs).

- **Vibe coding**, a term coined by Andrej Karpathy, involves deploying AI-generated code without thorough review, leading to significant failures, including a **6-hour outage** at Amazon that resulted in **6.3 million lost orders**.

- **Meta and Arm are collaborating to create the Arm AGI CPU**, a groundbreaking data center processor tailored for AI workloads, enhancing performance and efficiency beyond traditional CPUs.

- **TurboQuant** achieves **3.2× memory savings** by compressing model weights with **near-optimal distortion**, offering a **drop-in replacement for `nn.Linear`** that maintains performance.

- **The thermodynamics of computation** is fundamentally flawed due to its reliance on idealizations that ignore thermal fluctuations, leading to erroneous conclusions about non-dissipative processes at molecular scales.

- A **benchmark** was created to evaluate LLMs on **28 physics laws**, generating adversarial questions that exploit common cognitive biases and unit confusions, ensuring a rigorous assessment through **symbolic math** rather than subjective judgment.

- **PentaNet** introduces a **pentanary quantization** scheme {-2, -1, 0, +1, +2}, enhancing model capacity by **47%** compared to ternary quantization, while maintaining zero-multiplier inference benefits through simple bit-shifting.

- **Litellm versions 1.82.7 and 1.82.8 were compromised**, allowing a malicious .pth file to execute on every Python process start, scraping sensitive data like **SSH keys** and **API credentials** without requiring imports.

- The **first open-source implementation** of Hebbian fast-weight write-back for the BDH architecture enables the model to **rewrite its own decoder weights** during inference, utilizing sparse activation codes as addresses, which was previously unimplemented publicly.

- **New hafnium oxide memristors** could reduce AI energy consumption by **up to 70%** by mimicking the brain's efficient neural connections, enabling both storage and processing in a single location.

- The implementation of **TurboQuant** in Python introduces a novel approach to quantization by utilizing **random rotation** to achieve optimal 1D quantization without the need for calibration data or dataset-specific tuning, making it applicable across various contexts.

- **LVFace**, utilizing a **Vision Transformer (ViT)** backbone, reportedly outperforms **ArcFace** in facial recognition tasks, particularly in scenarios involving **partially occluded faces** like those wearing masks, as evidenced by its first-place finish in the **MFR-Ongoing challenge**.
