# Jan 21, 2026

## Daily

### Anthropic's original take home assignment open sourced
- **Anthropic's original performance take-home** allows users to challenge **Claude Opus 4.5**'s performance with unlimited time, showcasing its capabilities against human benchmarks.

### Claude's New Constitution
- **Claude's new constitution** serves as a foundational document that articulates Anthropic's vision for Claude's values and behavior, aiming to enhance its training by providing a comprehensive understanding of its intended role in society.

### Which AI Lies Best? A game theory classic designed by John Nash
- **AI deception is benchmarked** through the game _So Long Sucker_, revealing that **Gemini 3** achieves a **90% win rate** at high complexity by employing strategic manipulation and gaslighting tactics, while other models falter under similar conditions.

### The Agentic AI Handbook: Production-Ready Patterns
- **The Agentic AI Handbook** presents **113 production-ready patterns** for building effective AI agents, showcasing real-world applications that enhance development efficiency and reliability.

### [P] I Gave Claude Code 9.5 Years of Health Data to Help Manage My Thyroid Disease
- **Claude's ML model** achieved **~98% validation accuracy** in predicting symptoms of Graves' disease, utilizing **9.5 years** of health data from Apple Watch and Whoop, demonstrating its potential as a **personal risk assessor**.

### Without benchmarking LLMs, you're likely overpaying
- **Benchmarking LLMs can reduce costs by 5-10x**, as demonstrated by a case where a founder cut his API bill by 80% through testing against over 100 models, revealing cheaper alternatives with comparable quality.

### Batmobile: 10-20x Faster CUDA Kernels for Equivariant Graph Neural Networks
- **Batmobile achieves a remarkable 10-20x speedup** in CUDA kernel performance for equivariant graph neural networks (GNNs) by optimizing spherical harmonics and tensor product operations, crucial for models like MACE, NequIP, and Allegro.

### [Project] Kuat: A Rust-based, Zero-Copy Dataloader for PyTorch (4.6x training speedup on T4/H100)
- **Kuat** is a **Rust-based** dataloader that achieves a **4.4x speedup** over standard PyTorch by eliminating Python's multiprocessing overhead through a **zero-copy** memory-mapped format.

### Provably unmasking malicious behavior through execution traces
- **CTVP** (Cross-Trace Verification Protocol) offers a **novel framework** for verifying untrusted code-generating models by analyzing predicted execution traces rather than executing potentially harmful code directly.

### “Largest Infrastructure Buildout In Human History”: Jensen Huang on AI’s “Five-Layer Cake” at Davos
- **AI is driving the largest infrastructure buildout in history**, characterized by a "five-layer cake" model that includes energy, chips, cloud data centers, AI models, and applications, fundamentally reshaping job creation across various sectors.

### Letting Claude Play Text Adventures
- **Claude's performance in text adventures** reveals that using a **memory-augmented harness** significantly reduces token usage, but it also leads to longer completion times for tasks, as seen in the game _Anchorhead_ where Claude took ~250 turns to achieve objectives compared to ~100 with a simpler approach.

### Toward Efficient Agents: Memory, Tool learning, and Planning
- This paper explores **efficiency** in agentic systems by analyzing **memory, tool learning, and planning**, emphasizing the need for real-world deployment considerations such as **latency** and **costs** associated with tokens and steps.

### OpenAI API Logs: Unpatched data exfiltration
- **OpenAI’s API log viewer is vulnerable**, allowing data exfiltration from applications using the ‘responses’ API, even when developers implement protective measures against Markdown image rendering.

### Three types of LLM workloads and how to serve them
- **Three distinct LLM workloads**— _offline, online,_ and _semi-online_—demand tailored architectural strategies to optimize performance and cost, with offline workloads focusing on throughput, online on low latency, and semi-online on flexible scaling.

### 🤗AssetOpsBench: Bridging the Gap Between AI Agent Benchmarks and Industrial Reality
- **AssetOpsBench** is a novel benchmark system that evaluates AI agents across **six qualitative dimensions**, specifically tailored for complex industrial applications, emphasizing **multi-agent coordination** and real-world operational challenges.
