# Jul 11, 2026

## Daily

## Weekly

### Show HN: Getting GLM 5.2 running on my slow computer

- **colibrì** enables running the **GLM-5.2 (744B-parameter MoE)** model on consumer hardware with **~25 GB of RAM**, utilizing a unique architecture that streams experts from disk, allowing for efficient memory usage and processing.

### Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit

- The **MiMo-V2.5 series** leverages **Hybrid Sliding Window Attention (SWA)** to achieve a **7× reduction** in both compute and KVCache storage costs, enhancing efficiency for long-context and multimodal tasks.

### Computation as a universal and fundamental concept

- **Computation's limits** were first established by Alan Turing, who demonstrated that certain problems, like the halting problem, are unsolvable by any algorithm, highlighting the inherent boundaries of computational power.

### Silent speech with ultrasound

- **Model predicts speech from ultrasound recordings of the tongue**, achieving a **15.6% word error rate** on open-vocabulary speech, which is competitive with lip-reading methods trained on larger datasets.

### Nvidia, CoreWeave, and Nebius: Inside the Circular Financing of the GPU Boom

- **Neoclouds like CoreWeave and Nebius are capitalizing on the GPU boom by securing massive contracts, with commitments from Microsoft and Meta exceeding $120 billion, significantly outpacing their current revenues.**

### Show HN: Reame – a CPU inference server that gets faster as it runs

- **Reame is a pioneering LLM inference server** that optimizes for low-cost CPU hardware, achieving **100% accuracy** on long-context tasks with minimal resource expenditure, making it ideal for repetitive AI workloads on existing infrastructure.

### VultronRetriever family of models released on HuggingFace!

- The **VultronRetriever family of models** excels in performance, with **VultronRetrieverPrime-8B** ranking as the **global #1** on the MTEB Leaderboard, showcasing a **16x smaller index storage** and **12x higher throughput** than previous leaders.

### Towards Free Normalization: Fusing Normalization into GEMM and Attention Kernels

- **Novel kernel fusion techniques** for normalization operations like LayerNorm and RMSNorm can achieve up to **90% latency reduction** by integrating them with GEMM and attention kernels, significantly enhancing performance in deep learning models.
