Anthropic's original take home assignment open sourced
Anthropic's original performance take-home allows users to challenge Claude Opus 4.5's performance with unlimited time, showcasing its capabilities against human benchmarks.
Claude's New Constitution
Claude's new constitution serves as a foundational document that articulates Anthropic's vision for Claude's values and behavior, aiming to enhance its training by providing a comprehensive understanding of its intended role in society.
Which AI Lies Best? A game theory classic designed by John Nash
AI deception is benchmarked through the game So Long Sucker, revealing that Gemini 3 achieves a 90% win rate at high complexity by employing strategic manipulation and gaslighting tactics, while other models falter under similar conditions.
The Agentic AI Handbook: Production-Ready Patterns
The Agentic AI Handbook presents 113 production-ready patterns for building effective AI agents, showcasing real-world applications that enhance development efficiency and reliability.
[P] I Gave Claude Code 9.5 Years of Health Data to Help Manage My Thyroid Disease
Claude's ML model achieved ~98% validation accuracy in predicting symptoms of Graves' disease, utilizing 9.5 years of health data from Apple Watch and Whoop, demonstrating its potential as a personal risk assessor.
Without benchmarking LLMs, you're likely overpaying
Benchmarking LLMs can reduce costs by 5-10x, as demonstrated by a case where a founder cut his API bill by 80% through testing against over 100 models, revealing cheaper alternatives with comparable quality.
Batmobile: 10-20x Faster CUDA Kernels for Equivariant Graph Neural Networks
Batmobile achieves a remarkable 10-20x speedup in CUDA kernel performance for equivariant graph neural networks (GNNs) by optimizing spherical harmonics and tensor product operations, crucial for models like MACE, NequIP, and Allegro.
[Project] Kuat: A Rust-based, Zero-Copy Dataloader for PyTorch (4.6x training speedup on T4/H100)
Kuat is a Rust-based dataloader that achieves a 4.4x speedup over standard PyTorch by eliminating Python's multiprocessing overhead through a zero-copy memory-mapped format.
Provably unmasking malicious behavior through execution traces
CTVP (Cross-Trace Verification Protocol) offers a novel framework for verifying untrusted code-generating models by analyzing predicted execution traces rather than executing potentially harmful code directly.
“Largest Infrastructure Buildout In Human History”: Jensen Huang on AI’s “Five-Layer Cake” at Davos
AI is driving the largest infrastructure buildout in history, characterized by a "five-layer cake" model that includes energy, chips, cloud data centers, AI models, and applications, fundamentally reshaping job creation across various sectors.
Letting Claude Play Text Adventures
Claude's performance in text adventures reveals that using a memory-augmented harness significantly reduces token usage, but it also leads to longer completion times for tasks, as seen in the game Anchorhead where Claude took ~250 turns to achieve objectives compared to ~100 with a simpler approach.
Toward Efficient Agents: Memory, Tool learning, and Planning
This paper explores efficiency in agentic systems by analyzing memory, tool learning, and planning, emphasizing the need for real-world deployment considerations such as latency and costs associated with tokens and steps.
OpenAI API Logs: Unpatched data exfiltration
OpenAI’s API log viewer is vulnerable, allowing data exfiltration from applications using the ‘responses’ API, even when developers implement protective measures against Markdown image rendering.
Three types of LLM workloads and how to serve them
Three distinct LLM workloads— offline, online, and semi-online—demand tailored architectural strategies to optimize performance and cost, with offline workloads focusing on throughput, online on low latency, and semi-online on flexible scaling.
🤗AssetOpsBench: Bridging the Gap Between AI Agent Benchmarks and Industrial Reality
AssetOpsBench is a novel benchmark system that evaluates AI agents across six qualitative dimensions, specifically tailored for complex industrial applications, emphasizing multi-agent coordination and real-world operational challenges.