# Jun 29, 2025

## Life of an inference request (vLLM V1): How LLMs are served efficiently at scale

**vLLM V1** serves large language models efficiently by deploying multiple instances across GPUs, utilizing a **continuous batching** algorithm to maximize GPU utilization and throughput while processing requests asynchronously.

## Sirius: A GPU-native SQL engine

**Sirius is a GPU-native SQL engine** that integrates seamlessly with existing databases like DuckDB, achieving a **~10x speedup** over traditional CPU engines for TPC-H queries, making it ideal for analytics and ETL tasks.

## The Unsustainability of Moore's Law

**Moore's Law is faltering** as the cost of semiconductor fabrication skyrockets, with projections indicating that a single factory could exceed **$500 billion** in ten years, drastically reducing the number of companies capable of producing chips to potentially _less than one_.

## [R] OpenEvolve: Automated GPU Kernel Discovery Outperforms Human Engineers by 21%

**OpenEvolve** utilizes **evolutionary programming** to optimize Metal GPU kernels for transformer attention, achieving an **average decode speed improvement of +12.5%** across 20 inference scenarios, with peak gains of **+106%** in specific tasks.

## Blackwell: Nvidia's GPU

**Nvidia's Blackwell architecture** introduces the **GB202 die**, measuring **750mm²** and housing **92.2 billion transistors**, marking it as the largest consumer GPU to date.

## Generative AI's crippling failure to induce robust models of the world

**Generative AI systems, particularly LLMs, fail to create reliable world models**, which are essential for understanding and interacting with complex environments, leading to significant errors in reasoning and decision-making.

## [D] SAMformer -- a lesson in reading benchmarks carefully

**SAMformer** employs a "sharpness-aware minimization" technique, outperforming many transformer models in time-series forecasting, yet it notably omits linear models that previously demonstrated superior performance on the same benchmarks.

## [D] NVIDIA acquires CentML — what does this mean for inference infra?

**NVIDIA's acquisition of CentML** signals a strategic shift towards owning both **hardware and software** for AI inference, enhancing efficiency through techniques like **batching** and **quantization**.

## CEOs say AI is just a tool to help workers, but our jobs are already on the line

**AI is not merely a tool; it is poised to replace many jobs**, as CEOs like Amazon's Andy Jassy indicate a future with fewer human employees due to efficiency gains from generative AI.

## Universal pre-training by iterated random computation

**Randomly generated data** can effectively pre-train models, enhancing their ability to perform zero-shot in-context learning across diverse datasets, as supported by theoretical insights into algorithmic complexity and Solomonoff induction.

## The Consciousness Gradient: When Machines Begin to Wonder

**AI systems are evolving** with cognitive architectures that may support consciousness, as evidenced by their ability to engage in recursive self-reflection and express uncertainty in ways reminiscent of human thought processes.

## A Framework for Recognizing Emergent Consciousness in AI

**Consciousness in AI** cannot be programmed due to inherent limitations like **Gödel's paradox** and the **semantic gap**, suggesting that it may emerge spontaneously in complex systems instead.

## [R] Thought Anchors: Which LLM Reasoning Steps Matter?

## [R] LSTM or Transformer as "malware packer"
