# Jan 20, 2026

### The coming industrialisation of exploit generation with LLMs

- **LLMs like GPT-5.2 and Opus 4.5 can autonomously generate over 40 distinct exploits for a zeroday vulnerability, demonstrating a shift towards the _industrialisation_ of exploit generation in cybersecurity.**

### X For You Feed Algorithm

- The **For You feed algorithm** on X utilizes a **Grok-based transformer model** to seamlessly integrate in-network and out-of-network content, predicting engagement probabilities without relying on hand-engineered features.

### The assistant axis: situating and stabilizing the character of LLMs

- The **Assistant Axis** is a newly defined direction in the persona space of large language models, linking **Assistant-like behavior** to specific neural activity patterns that can be monitored and stabilized to prevent harmful persona drift.

### Scaling long-running autonomous coding

- **Cursor's autonomous coding agents** successfully generated over **1 million lines of code** while attempting to build a web browser from scratch, showcasing the potential of AI in software development.

### Bypassing Gemma and Qwen safety with raw strings

- The **apply_chat_template()** function is crucial for maintaining safety alignment in open-source LLMs, as omitting it can lead to models generating harmful content, revealing that safety is not inherent in the model weights but rather in the formatting of prompts.

### I Gave Claude Code 9.5 Years of Health Data to Help Manage My Thyroid Disease

- **Claude's ML model** achieved **~98% validation accuracy** in predicting symptoms of Graves' disease, utilizing **9.5 years** of health data from Apple Watch and Whoop, demonstrating its potential as a **personal risk assessor**.

### Without benchmarking LLMs, you're likely overpaying 5-10x

- **Benchmarking LLMs can reduce costs by 5-10x**, as demonstrated by a case where a founder cut his API bill by 80% through testing against over 100 models, revealing cheaper alternatives with comparable quality.

### Weight Transfer for RL Post-Training in under 2 seconds

- Achieving **1.3-second** cross-machine parameter updates for Kimi-K2 (1T parameters) demonstrates a significant advancement in weight transfer efficiency, utilizing **256 training GPUs** (BF16) and **128 inference GPUs** (FP8) through **RDMA point-to-point communication**.

### tested file based memory vs embedding search for my chatbot. the difference in retrieval accuracy was bigger than i expected

- **File-based memory** outperformed **embedding search** in retrieval accuracy, especially for complex queries, achieving up to **75% accuracy** on temporal queries compared to **40%** for embedding search.

### Differential Transformer V2

- **Differential Transformer V2 (DIFF V2)** enhances inference efficiency and training stability for large language models by introducing additional parameters for faster decoding and eliminating the need for custom attention kernels, thus matching the baseline Transformer’s speed.

### Cisco and OpenAI redefine enterprise engineering with AI agents

### Our approach to age prediction

### Multimodal reinforcement learning with agentic verifier for AI agents

- **Argos** is a novel verification framework that enhances **multimodal reinforcement learning** by rewarding AI agents for producing correct answers grounded in visual and temporal evidence, thus improving reliability and reducing errors in real-world applications.

### Kinematic Fingerprints: Predicting sim-to-real transfer success from movement signatures

- **Kinematic fingerprints** predict the success of sim-to-real transfer by analyzing movement signatures, achieving **85-90% accuracy** on unseen policies across various robot platforms.
