# Jul 12, 2024

- **FlashAttention-3** introduces techniques exploiting **asynchrony and low-precision (FP8)** on Hopper GPUs, achieving **1.5-2.0x speed improvements** and **75% utilization** of H100 GPU capabilities, significantly enhancing the efficiency of attention mechanisms in large language models (LLMs).

- The **Physics-based Deep Learning Book** offers a **comprehensive guide** to integrating deep learning with physical simulations, featuring **hands-on Jupyter notebook examples** for immediate application.

- **Korvus integrates the entire RAG pipeline into a single Postgres database query**, offering a unified search SDK with Python, JavaScript, and Rust bindings for high-performance, customizable search functionalities.

- **AuraFlow v0.1** is introduced as a **large, open-source, flow-based generation model** capable of text-to-image generation, challenging the notion that open-source AI development has stalled.

- **Free-threaded CPython**, an experimental feature in **CPython 3.13**, enables multiple threads to run in parallel without the Global Interpreter Lock (GIL), aiming to improve multi-threaded performance by utilizing multiple CPU cores more effectively, as detailed in [PEP 703](https://peps.python.org/pep-0703/).

- **`mandala` streamlines ML experiment tracking** by capturing inputs, outputs, and code dependencies with the `@op` decorator, ensuring no computation is repeated unnecessarily.

- **StreamVC** is a **real-time voice conversion technology** that maintains the original speech's content and prosody while adopting the voice timbre of any target speech, designed for **low-latency performance** on mobile platforms.

- EvolutionaryScale, founded by former Meta scientists, has **unveiled a protein language model, ESM3**, which has successfully **engineered new fluorescent proteins**, marking a significant advancement in AI-driven biological design.

- **WildGaussians** introduces a **novel approach** to **3D Gaussian Splatting (3DGS)**, enhancing its capability to handle **occlusions and appearance changes** in dynamic, in-the-wild scenes.

- **Cradle** introduces a **General Computer Control (GCC)** setting, using **screenshots for input** and **keyboard/mouse actions for output**, to enable foundation agents to interact with any software, overcoming the challenge of environment encapsulation differences.

- **Memory^3** introduces a **novel approach** to language modeling by equipping large language models (LLMs) with **explicit memory**, significantly reducing parameter size, training, and inference costs while outperforming larger models and text retrieval-augmented generation (RAG) models. [Read the paper](https://arxiv.org/pdf/2407.01178)

- The **dendristor model** integrates synaptic organization with dendritic tree-like morphology, utilizing multigate silicon nanowire transistors for **neuromorphic computation**, emulating **visual motion perception** akin to that in the retina.

- **Self-supervised learning** with **DINOv2 ViT weights** from Facebook Research enables **rich image segmentation** by finetuning with Low-Rank Adaptation (LoRA) and simple decoders, achieving **solid validation IoU scores**.

- **FlashAttention-3** introduces techniques on Hopper GPUs for **1.5-2.0x speed** over FlashAttention-2, achieving up to **740 TFLOPS** with FP16 and nearly **1.2 PFLOPS** with FP8, reducing error by **2.6x** in low-precision contexts. [Paper](https://tridao.me/publications/flash3/flash3.pdf)

- **Google Research** introduces a method to **revolutionize online handwriting recognition** by leveraging **large vision-language models (VLMs)**, which utilize new representations and tokenizers for improved accuracy.
