# Aug 23, 2025

## Writing Speed-of-Light Flash Attention for 5090 in CUDA C++
- **Implementing Flash Attention for 5090 in CUDA C++** reveals that the author navigated the limitations of Triton by leveraging CUDA C++ to optimize attention mechanisms, achieving significant performance improvements over existing implementations.

## 450× Faster Joins with Index Condition Pushdown
- **New optimization using Index Condition Pushdown (ICP) significantly enhances straddled joins in Readyset**, allowing for efficient retrieval of only necessary rows and reducing unnecessary data reads during cache misses.

## Launch HN: BlankBio (YC S25) - Making RNA Programmable
- **BlankBio** is developing a **computational toolkit** for mRNA design, enabling biologists to create effective therapeutic sequences through **RNA foundation models** trained on unlabeled data, which significantly reduces reliance on noisy experimental data.

## Why do BYOL/JEPA-like models work? How does EMA prevent model collapse?
- **BYOL and JEPA models** excel in learning **semantic embeddings** without direct target reconstruction, leveraging **Exponential Moving Average (EMA)** to prevent model collapse during training.

## RIKEN, Japan’s Leading Science Institute, Taps Fujitsu and NVIDIA for Next Flagship Supercomputer
- **FugakuNEXT**, Japan's next flagship supercomputer, will integrate **Fujitsu** and **NVIDIA** technologies to address critical scientific challenges, emphasizing a collaborative design approach for enhanced performance and innovation.

## Show HN: OctaneDB – Fast, Open-Source Vector Database for Python
- **OctaneDB** delivers **10x faster** performance than competitors like Pinecone and ChromaDB, making it ideal for AI/ML applications that demand rapid similarity searches.

## Robots can now learn to use tools just by watching us
- **Robots can learn tool use by observing human actions in videos**, utilizing a framework called **Tool-as-Interface**, which allows them to adapt skills without extensive programming or specialized equipment.

## Accelerating life sciences research
- **NVIDIA's innovations** at the upcoming Hot Chips conference will showcase how **NVLink, Spectrum-X Ethernet, Blackwell architecture, and CUDA** are revolutionizing AI inference across global data centers, enhancing performance and efficiency.

## Relational PDF Recall (RFC + PoC) – Structured storage + overlay indexing experiment
- **Relational database structures inside PDFs** can enhance AI recall by enabling **channel splitting** and **relational indexing**, as demonstrated in the draft RFC and PoC.

## I built a ML-regression model for Biathlon that beats current betting market odds
- The **ML-regression model** for biathlon predicts outcomes with a **MAE of 0.14** and an **R² of ~62%**, significantly outperforming current betting market odds and reducing random guessing error by nearly half.

## DRAMA Model Inference Efficiency Boosted by 1.7x-2.3x
- The **DRAMA model** achieves a **1.7x-2.3x boost** in inference efficiency through the use of **Nested Jagged Tensors (NJT)**, enhancing its viability for production environments with variable-length sequences.
