# Aug 3, 2025

Daily

Weekly

- **Persona vectors** are patterns of neural activity that help monitor and control the **character traits** of language models, enabling developers to understand and mitigate undesirable personality shifts during training and deployment.

- **Token costs are rising**, contradicting the expectation that cheaper models would improve consumer AI margins, as evidenced by companies like Windsurf and Claude Code facing severe financial strain despite lower LLM inference costs.

- **Micron's new SSD trio** features the **9650**, a PCIe Gen 6 drive optimized for speed with up to **5.5 million IOPS** and **28 GBps** read speeds, the **6600 ION** offering high capacity up to **122.88 TB** with QLC flash, and the **7600** designed for low latency in capacities up to **15.36 TB**.

- **SuperVM**, a specialized bytecode optimizer, achieves an impressive **99.8 FPS**, outperforming Copilot models by **2X** in a simple fractal generation task, demonstrating the effectiveness of deterministic systems over statistical ones.

- The **GPT with Modified Memorizing Transformer** enhances traditional transformer models by integrating a memorization mechanism, allowing for improved retention of information over longer contexts.

- **Kimi K2** is a **Mixture-of-Experts (MoE)** model with **32 billion activated parameters** and **1 trillion total parameters**, utilizing the **MuonClip optimizer** to enhance training stability and token efficiency, pre-trained on **15.5 trillion tokens** without loss spikes.

- **Current LLMs are reaching their limits** as they can only replicate existing human knowledge, necessitating a shift towards **experiential learning** where AI interacts with the real world to discover new principles and overcome inherited biases.

- **Reflecting on 8 years of model inference** reveals significant advancements in **tooling** that enhance efficiency and accuracy in machine learning applications, as discussed in the [Lightning blog](https://lightning.ai/blog/evolution-of-model-inference).

- The **Mamba Adaptive Anomaly Transformer (MAAT)** enhances the **Anomaly Transformer** framework by integrating a **Mamba-selective state-space model** and **sparse attention**, achieving superior performance in anomaly detection for time series data.
