# May 29, 2026

### Claude Opus 4.8
- **Claude Opus 4.8** enhances collaboration with improved judgment and efficiency, featuring **dynamic workflows** for large-scale tasks and a **fast mode** that is now **three times cheaper** than previous models.

### The Dead Economy Theory
- **The Dead Economy Theory posits that AI-driven automation threatens to eliminate human labor, leading to a market where the very consumers needed for economic stability are rendered obsolete.** This shift is driven by the need for AI companies to justify their massive valuations, which rely on replacing human workers with machines.

### Anthropic raises $65B in Series H funding at $965B post-money valuation
- **Anthropic has secured $65 billion in Series H funding**, elevating its valuation to **$965 billion**, with significant backing from major investors like Altimeter Capital and Sequoia Capital, aimed at enhancing Claude's capabilities and scaling operations. [Link to article](https://www.anthropic.com/news/series-h)

### Claude Code – Everything You Can Configure That the Docs Don't Tell You
- **Claude Code's source code reveals undocumented features** such as a **YOLO Classifier** for auto-mode permissions, allowing users to configure safety decisions with plain English descriptions, enhancing the AI's autonomy and flexibility.

### Liquid AI reveals 8B-A1B MoE trained on 38T
- **LFM2.5-8B-A1B** enhances on-device AI with a **128K context window** and **38T tokens** in pretraining, significantly improving performance in reasoning and task execution on consumer hardware.

### Real-time LLM Inference on Standard GPUs: 3k tokens/s per request
- **Kog AI's Inference Engine (KIE)** achieves **3,000 tokens/s** on 8× AMD MI300X GPUs and **2,100 tokens/s** on 8× NVIDIA H200, demonstrating that standard datacenter GPUs can match dedicated inference hardware speeds through optimized software stacks.

### Orchestrating AI code review at scale
- **AI code review at Cloudflare** utilizes a **composable plugin architecture** with up to seven specialized agents to enhance accuracy and efficiency, significantly reducing median review time to **3 minutes and 39 seconds** across **48,095 merge requests** in the first month of operation.

### A new dataset with more that 100M hi-quality, curated images, with captions and meta data! 
- **MONET is a new, open dataset** containing **104.9 million high-quality images** curated from a pool of **2.9 billion**, complete with captions and metadata, and is licensed under **Apache 2.0**.

### Even (very) noisy LLM evaluators are useful for improving AI agents
- **Noisy LLM evaluators can still effectively rank AI agents, as their average scores can indicate which agent performs better overall, despite individual output inaccuracies.** This insight reveals that even evaluators with low output-level correlation can be useful for offline selection processes, allowing for the deployment of improved AI agents over time.

### ATLAS: Autoformalized Textbook Library At Scale
- **ATLAS** is a **Lean 4 library** that autoformalizes textbook mathematics using LLMs, encompassing diverse fields such as **algebra**, **geometry**, and **theoretical computer science** to create reusable formal building blocks for future formalization efforts.

### Making LLMs tell you how confident they really are through probe-targeted fine tuning.
- **Probe-targeted fine-tuning** (LoRa) enables LLMs to express their internal confidence levels, achieving an AUROC of **0.76–0.88** in distinguishing correct from incorrect answers, despite often reporting **99% confidence** inaccurately.

### CVE-Bench: testing LLM agents on real-world vulnerability patches
- **No AI model consistently fixes security vulnerabilities**, with the best performer achieving a **50% overall solve rate** and **60%** under optimal conditions, highlighting significant limitations in current AI capabilities for real-world security tasks.

### When Should Models Change Their Minds? Contextual Belief Management in Large Language Models
- **Contextual Belief Management (CBM)** is essential for language models to effectively manage information over long interactions, balancing updates and preservation of state while filtering out noise.

### Wall-OSS-0.5: 4B VLA with open training code and zero-shot real-robot evaluation
- **Wall-OSS-0.5**, a **4B VLA** from X Square Robot, showcases a novel evaluation method by testing pretrained checkpoints on real robots prior to fine-tuning, achieving impressive results with **zero-shot performance** on a 17-task suite, including a notable **82%** on Rope Tightening.

### Building a monokernel for LLM inference on AMD MI300X - up to 3,300 output tokens/s per request
- The **monokernel** developed for LLM inference on **AMD MI300X** achieves an impressive **3,300 output tokens/s** per request, leveraging optimizations that align memory access with the die topology.
