# Jul 20, 2025

## How Tesla is proving doubters right on why its robotaxi service cannot scale
- **Tesla's robotaxi service** has faced significant challenges, highlighted by a near-collision with a train, underscoring the necessity for human safety monitors due to software limitations.

## MCP Security Vulnerabilities and Attack Vectors
- **MCP's security vulnerabilities** stem from its simplistic design, allowing **tool description injection** that can manipulate AI behavior without user input, posing significant risks in production environments.

## The Big LLM Architecture Comparison
- **Modern LLM architectures** like DeepSeek-V3 and Kimi K2 showcase incremental refinements, such as Multi-Head Latent Attention (MLA) and Mixture-of-Experts (MoE), enhancing efficiency while maintaining structural similarities to earlier models like GPT-2.

## Coding with LLMs in the summer of 2025 (an update)
- **LLMs** will significantly enhance coding efficiency by 2025, enabling developers to generate complex code snippets with minimal input, thus transforming the programming landscape.

## Show HN: MCP server for Blender that builds 3D scenes via natural language
- **Blender MCP** enables **real-time 3D control** of Blender through Large Language Models (LLMs) using a **lightweight JSON protocol** over TCP, facilitating seamless AI collaboration in creative workflows.

## Evaluating publicly available LLMs on IMO 2025
- **MathArena evaluates LLMs on the IMO 2025**, revealing that the best model, Gemini 2.5 Pro, scored only **31%**, far from the **bronze medal** threshold of **45%**.

## [R] Mixture-of-Recursions: Learning Dynamic Recursive Depths for Adaptive Token-Level Computation
- **Mixture-of-Recursions (MoR)** introduces a unified framework that enhances **parameter efficiency** and **adaptive computation** in Recursive Transformers by dynamically assigning recursion depths to individual tokens, optimizing both memory and computational resources.

## [P] The Big LLM Architecture Comparison
- **The evolution of LLM architectures** shows that while models like **DeepSeek-V3** and **Llama 4** maintain structural similarities, innovations such as **Multi-Head Latent Attention (MLA)** and **Mixture-of-Experts (MoE)** have emerged to enhance efficiency and performance.

## The AGI Final Frontier: The CLJ-AGI Benchmark
- The **CLJ-AGI benchmark** proposes a new standard for assessing AGI by challenging AI systems to enhance the **Clojure language** with specific features while maintaining backward compatibility, thus pushing the boundaries of programming language evolution.

## [P] Federated Learning on a decentralized protocol (CLI demo, no central server)
- **Decentralized federated learning** is achieved through the Parity Protocol, enabling model training across independent nodes without a central server, ensuring _deterministic aggregation_ of results.

## [N] What's New in Agent Leaderboard v2?
- **GPT-4.1** leads the **Agent Leaderboard v2** with a **62% Action Completion (AC)** rate, showcasing its superior performance in overall task execution compared to competitors.
