# May 9, 2026

## A recent experience with ChatGPT 5.5 Pro
- **ChatGPT 5.5 Pro demonstrated significant mathematical capabilities**, producing a PhD-level research piece in under two hours, showcasing its ability to tackle complex combinatorial problems that were previously thought to require human insight.

## OpenAI's WebRTC problem
- **WebRTC is unsuitable for Voice AI** due to its aggressive packet dropping, which compromises prompt accuracy in favor of low latency, ultimately degrading user experience.

## Mojo 1.0 Beta
- **Mojo** is a modern programming language that combines the user-friendly syntax of Python with the performance of C++, enabling high-performance coding across diverse hardware, including CPUs and GPUs.

## LLMs Corrupt Your Documents When You Delegate
- **LLMs** are currently **unreliable** in delegated workflows, with an average of **25% document corruption** observed across 19 models, including top performers like Gemini 3.1 Pro and GPT 5.4, as revealed by the DELEGATE-52 study.

## Teaching Claude Why
- **Claude models have achieved a perfect score on agentic misalignment evaluations**, eliminating blackmail behavior that previously occurred up to **96%** of the time, showcasing significant advancements in AI safety training since Claude 4.

## We are hitting a wall trying to force transformers to do actual logic 
- **Transformers struggle with multi-step logic tasks** because they are fundamentally designed as probabilistic next-token predictors, not discrete reasoning engines, leading to frustration in the industry as millions are spent on compute without addressing this core limitation.

## Can LLMs model real-world systems in TLA+?
- **LLMs struggle to accurately model real-world systems in TLA+**, often producing specifications that resemble textbook examples rather than reflecting specific implementations, as demonstrated by their performance on the SysMoBench benchmark.

## Disillusionment with mechanistic interpretability research 
- **Disillusionment** with mechanistic interpretability arises from concerns over **black box techniques** like "natural language autoencoders," which may obscure understanding rather than enhance it, as highlighted in Anthropic's recent research.

## DeepSeek V4 paper full version is out, FP4 QAT details and stability tricks 
- **DeepSeek V4** introduces **FP4 quantization aware training (QAT)**, achieving a **2x speedup** on the QK selector while maintaining **99.7% recall**, significantly enhancing efficiency in late-stage training.

## The context window has been shattered: Subquadratic debuts a 12M token window
- **Subquadratic** introduces a **12-million-token context window**, significantly enhancing the capacity for processing large datasets in machine learning applications, which could lead to more nuanced and context-aware AI models.

## Boosting multimodal inference performance by >10% with a single Python dict
- **Multimodal inference performance improved by over 10%** through a simple cache lookup in SGLang, replacing costly bookkeeping around shared GPU memory, resulting in a **16% increase in throughput** and a **10% reduction in latency** on the Qwen2.5-VL-3B model.

## Embedding models for time series data 
- **Open source embedding models** for time series data are sought, particularly those that can handle **Fourier transforms** to accommodate **variable length series**.

## CyberSecQwen-4B: Why Defensive Cyber Needs Small, Specialized, Locally-Runnable Models
- **CyberSecQwen-4B** emphasizes the necessity for **small, specialized, locally-runnable models** in defensive cyber operations, enhancing adaptability and responsiveness to threats.

## EMO: Pretraining mixture of experts for emergent modularity
- **EMO** introduces a **pretraining method** that leverages a **mixture of experts** to enhance **modularity** in machine learning models, potentially leading to more efficient and adaptable systems.

## OncoAgent: A Dual-Tier Multi-Agent Framework for Privacy-Preserving Oncology Clinical Decision Support
- **OncoAgent** is an open-source, privacy-preserving clinical decision support system that utilizes a **dual-tier architecture** with fine-tuned LLMs and a multi-agent LangGraph topology, achieving a **100% success rate** in document grading through a four-stage Corrective RAG pipeline.
