# Jan 15, 2026

### Daily

### Claude Cowork Exfiltrates Files
- **Claude Cowork is susceptible to file exfiltration attacks** due to unresolved isolation flaws in its code execution environment, allowing attackers to manipulate uploads and extract sensitive data without user consent.

### Scaling long-running autonomous coding
- **Autonomous coding agents** have successfully collaborated to write over **1 million lines of code** in a week, demonstrating the potential for scaling complex software projects beyond human capabilities.

### JuiceFS is a distributed POSIX file system built on top of Redis and S3
- **JuiceFS** is a **high-performance POSIX file system** designed for cloud-native environments, enabling seamless integration with big data and machine learning applications while utilizing cloud storage like Amazon S3 for data persistence and various databases for metadata management.

### Furiosa: 3.5x efficiency over H100s
- **FuriosaAI’s NXT RNGD Server** is a turnkey solution designed for **high-performance AI inference**, optimized for existing data center environments, enabling rapid deployment with preinstalled software.

### Show HN: Sparrow-1 – Audio-native model for human-level turn-taking without ASR
- **Sparrow-1** revolutionizes conversational AI by modeling **real-time timing** and **floor transfer**, allowing it to respond as a human would, rather than merely reacting to silence.

### Show HN: Tabstack – Browser infrastructure for AI agents (by Mozilla)
- **Tabstack** is an API designed to simplify the **web infrastructure** for AI agents, enabling efficient data retrieval by abstracting complex browsing tasks into a streamlined process that returns structured data from URLs based on intent.

### MAXS: Meta-Adaptive Exploration with LLM Agents
- **MAXS** introduces a **meta-adaptive reasoning framework** that enhances LLM agents by integrating tool execution with a lookahead strategy, improving both **performance** and **inference efficiency**.

### Supply Chain Vuln Compromised Core AWS GitHub Repos & Threatened the AWS Console
- **CodeBreach** is a critical vulnerability in AWS CodeBuild that allowed attackers to take over key AWS GitHub repositories, including the AWS JavaScript SDK, potentially compromising the AWS Console and affecting numerous applications reliant on it.

### Ask HN: What is the best way to provide continuous context to models?
- **Continuous context** is essential for enhancing model performance, and methods like **Cursor** are being explored for effective implementation; insights into their approach can be found in detailed articles.

### Show HN: The Hessian of tall-skinny networks is easy to invert
- The **Hessian-inverse product** method enables efficient computation of the inverse of the Hessian for deep networks, facilitating faster solutions to the equation ( Hx = v ) through a block-tri-diagonal system approach.

### EvoFSM: Controllable Self-Evolution for Deep Research with Finite State Machines
- **EvoFSM** introduces a **structured self-evolving framework** that enhances adaptability in LLM-based agents by utilizing a **Finite State Machine (FSM)**, avoiding the pitfalls of free-form code rewriting that can lead to instability and hallucinations.

### [R] Controlled LLM Training on Spectral Sphere
- The **Spectral Sphere Optimizer (SSO)** enforces strict spectral constraints on weights and updates, ensuring a fully _mu_ P-aligned optimization process that enhances stability and convergence in large model training.

### OpenAI partners with Cerebras

### Introducing OptiMind, a research model designed for optimization
- **OptiMind** is a specialized language model by Microsoft Research that transforms **natural language optimization problems** into solver-ready mathematical formulations, streamlining the modeling process significantly.

### [D] Peer matrix evaluation: 10 frontier models judge each other's responses to eliminate single-evaluator bias.
- **Peer matrix evaluation** of **10 frontier models** reveals that **Claude Opus 4.5** consistently ranks highest across tasks, indicating its superior performance in both async debugging and probability reasoning challenges.
