# Aug 30, 2025

## Deploying DeepSeek on 96 H100 GPUs

- **DeepSeek's implementation on 96 H100 GPUs achieves a remarkable throughput of 52.3k input tokens and 22.3k output tokens per second per node**, utilizing prefill-decode disaggregation and large-scale expert parallelism, marking a significant advancement in open-source LLM performance.

## Hermes 4

- **Hermes 4** introduces a new class of **hybrid reasoning models** that enhance user interaction by allowing for longer, thoughtful responses and a more engaging, humanistic experience, free from corporate biases.

## AI models need a virtual machine

- **AI models require a standardized control layer akin to a virtual machine** to ensure security, isolation, and extensibility, enabling seamless integration into diverse software ecosystems while maintaining robust governance.

## From Multi-Head to Latent Attention: The Evolution of Attention Mechanisms

- **Attention mechanisms** have evolved from **Multi-Head Attention (MHA)** to **Multi-Head Latent Attention (MHLA)**, enhancing efficiency in processing context while maintaining performance in NLP tasks.

## The Theoretical Limitations of Embedding-Based Retrieval

- **Vector embeddings face inherent theoretical limitations** in retrieval tasks, which persist even with simple queries, challenging the assumption that better training data and larger models can overcome these issues.

## Data engineering and software engineering are converging

- **Data engineering and software engineering are merging**, necessitating a robust data infrastructure that enhances developer experience (DX) for real-time analytics and AI features, as exemplified by the open-source [MooseStack](https://github.com/514-labs/moosestack) toolkit for ClickHouse.

## SynthID

- **SynthID** is a watermarking tool that embeds imperceptible digital markers in AI-generated content, enhancing **transparency** and **trust** in generative AI outputs.

##  Scaling Inference: Lessons from Running Multiple Foundation Models in Production

- **Inference optimization** is critical when deploying multiple foundation models like LLaMA and Mistral, as **batching** can enhance cost-efficiency but may severely impact latency for interactive applications.
