# Jun 6, 2025

## Eleven v3
- **Eleven v3 (alpha)** introduces a **highly expressive Text to Speech model** that allows for dynamic conversations and emotional depth through audio tags, enhancing user engagement and realism in generated speech.

## Tokasaurus: An LLM Inference Engine for High-Throughput Workloads
- **Tokasaurus** is a new **LLM inference engine** designed for **high-throughput workloads**, achieving up to **3x** performance improvement over competitors like vLLM and SGLang through optimized CPU usage and advanced parallelism techniques.

## [R] LLMs are Locally Linear Mappings: Qwen 3, Gemma 3 and Llama 3 can be converted to exactly equivalent locally linear systems for interpretability
- **LLMs can be transformed into locally linear systems** that accurately reconstruct next-token outputs without altering model weights, enhancing interpretability through linear algebra techniques.

## Sandia turns on brain-like storage-free supercomputer – Blocks and Files
- **Sandia National Labs** has activated its **SpiNNaker 2 supercomputer**, a brain-inspired system that operates without traditional storage, mimicking **150 to 180 million neurons** to enhance computational efficiency for national security applications.

## Aurora, a foundation model for the Earth system
- **Microsoft's Aurora model** can deliver accurate **10-day weather forecasts** rapidly, showcasing its potential to transform not just weather predictions but also other Earth systems like air pollution and wave height.

## [R] Atlas: Learning to Optimally Memorize the Context at Test Time
- **ATLAS** introduces a **long-term memory module** that optimally memorizes context by leveraging both current and past tokens, significantly enhancing performance in autoregressive language modeling tasks.

## Free Gaussian Primitives at Anytime Anywhere for Dynamic Scene Reconstruction
- **FreeTimeGS** introduces a **novel 4D representation** of Gaussian primitives, enhancing the reconstruction of dynamic 3D scenes by allowing primitives to appear at any time and location, thus improving flexibility and reducing temporal redundancy.

## Machine Learning: The Native Language of Biology
- **Machine learning methods excel in modeling biological systems**, revealing complex, non-linear relationships that traditional mathematical approaches often overlook, thus suggesting a new paradigm for understanding biology.

## Switching to Mojo gave a 14% improvement over CUDA
- The **highly efficient matrix transpose kernel** implemented in Mojo achieves a remarkable bandwidth of **2775.49 GB/s**, demonstrating that Mojo can match **CUDA** performance on the same task, with optimizations leading to significant speed improvements.

## [R] What do you all think of the latest Apple paper on current LLM capabilities?
- The **latest Apple paper** reveals that **LLMs** (Large Language Models) and **LRMs** (Language Reasoning Models) exhibit **limited true reasoning capabilities**, particularly struggling with complex tasks that require human-like reasoning.

## Fuzzer Blind Spots (Meet Jepsen)
- **Jepsen uncovered a critical correctness bug** in TigerBeetle's query engine, despite extensive fuzz testing, highlighting the limitations of current testing methodologies in identifying issues in complex systems.

## Series C and Scale (Cursor)
- **Cursor has secured $900 million** in Series C funding, elevating its valuation to **$9.9 billion**, with backing from prominent investors like Thrive, Accel, and Andreessen Horowitz.

## [R] 100M Open source notebooklm speech model
- A **100M open source notebooklm speech model** has been developed using two **NVIDIA 4090 GPUs**, showcasing significant advancements in speech processing capabilities.

## [R] Better quantization: Yet Another Quantization Algorithm
- **Yet Another Quantization Algorithm (YAQA)** significantly enhances model output preservation post-quantization, achieving a **KL reduction of over 30%** compared to QTIP and outperforming Google's QAT model on Gemma 3.

## Workhorse LLMs: Why Open Source Models Dominate Closed Source for Batch Tasks
- **Open source LLMs** are increasingly outperforming closed source models like GPT-4o-mini and Gemini 2.5 Flash for common tasks, offering **significant cost savings** and improved performance, especially in batch processing scenarios.
