# Mar 8, 2026

## Daily Updates

- **How to run Qwen 3.5 locally**  
  **Qwen3.5** offers a range of models, including **35B**, **27B**, and **397B**, optimized for local deployment with **256K context** support across **201 languages**, excelling in tasks like coding and chat.

- **Cloud VM benchmarks 2026: performance/price for 44 VM types over 7 providers**  
  **AMD EPYC Turin** emerges as the **top performer** in the 2026 cloud VM benchmarks, significantly outperforming previous generations and offering exceptional value across various providers.

- **Files are the interface humans and agents interact with**  
  **Filesystems are gaining traction in AI** as they provide a persistent context layer that enhances agent functionality, allowing for more efficient coding and project management without the complexity of traditional databases.

- **Autoresearch: Agents researching on single-GPU nanochat training automatically**  
  **Autonomous AI agents** now conduct research by modifying code and training models without human intervention, leading to potentially faster advancements in AI capabilities.

- **SWE-CI: Evaluating Agent Capabilities in Maintaining Codebases via CI**  
  **SWE-CI** introduces a novel benchmark that evaluates agent capabilities in maintaining codebases through **Continuous Integration**, focusing on long-term _maintainability_ rather than just short-term functional correctness.

- **I don't know if my job will still exist in ten years**  
  **The future of software engineering is uncertain**, with the potential for AI to significantly alter job roles, possibly leading to a shift towards supervising AI agents rather than traditional coding tasks.

- **Phi-4-reasoning-vision and the lessons of training a multimodal reasoning model**  
  **Phi-4-reasoning-vision-15B** is a **15 billion parameter** multimodal reasoning model that excels in **math and science reasoning**, offering efficient performance across various vision-language tasks while minimizing compute costs.

- **NanoJudge: Instead of prompting a big LLM once, it prompts a tiny LLM thousands of times**  
  **NanoJudge** revolutionizes ranking by utilizing **thousands of pairwise comparisons** with tiny LLMs, yielding a mathematically rigorous leaderboard that overcomes traditional LLM limitations.

- **Graph-Oriented Generation (GOG): Replacing Vector R.A.G. for Codebases with Deterministic AST Traversal (70% Average Token Reduction)**  
  **Graph-Oriented Generation (GOG)** replaces Vector RAG by utilizing a **deterministic Symbolic Reasoning Model (SRM)** to parse codebases as a **Directed Acyclic Graph (DAG)**, achieving a **70% average token reduction** in processing.

- **Combining Stanford's ACE paper with the Reflective Language Model pattern - agents that write code to analyze their own execution traces at scale**  
  **Combining Stanford's ACE and Reflective Language Model (RLM) patterns enables agents to analyze execution traces more effectively** by utilizing a Recursive Reflector that programmatically queries data, overcoming limitations of single-pass analysis.

- **Introducing NNsight v0.6: Open-source Interpretability Toolkit for LLMs**  
  **NNsight 0.6** introduces significant enhancements, including **remote execution** capabilities via the National Deep Inference Fabric (NDIF), allowing users to run analyses without local GPU resources by simply adding `remote=True`.

- **We Turned Our Wireshark Wizard into a Markdown File**  
  **Rocky AI** has evolved from a simple chat widget to a sophisticated **Root Cause Analysis Agent**, leveraging human expertise codified into markdown files for effective troubleshooting of network issues.

- **We analyzed 4,000 Ethereum contracts by combining an LLM and symbolic execution and found 5,783 issues**  
  **SymGPT** effectively combines **large language models** and **symbolic execution** to identify **5,783 ERC rule violations** in **4,000 Ethereum contracts**, highlighting its capability to detect vulnerabilities, including **1,375** with potential for **financial theft**.

- **VeridisQuo - open-source deepfake detector that combines spatial + frequency analysis and shows you where the face was manipulated**  
  **VeridisQuo** introduces a method to detect deepfakes by combining **spatial and frequency** analysis, making it easier to identify manipulations in media.
