ML Times
Jan 14, 2026
Claude Cowork is susceptible to file exfiltration attacks due to unresolved isolation flaws in its code execution environment, allowing attackers to manipulate uploads and extract sensitive data without user consent.
vLLM's migration to the V1 engine has enabled a remarkable throughput of 2.2k tokens/s per H200 GPU, showcasing the effectiveness of optimizations like async scheduling and dual-batch overlap.
Dicer is now open-sourced, enabling developers to build fast, scalable, and highly available sharded services, addressing the limitations of traditional stateless and static sharding architectures.
Harmonic's pbcc is a custom Protocol Buffers compiler for Python, designed to enhance performance by generating specialized C++ code, thus enabling efficient handling of large datasets with a cleaner API.
Systematic testing using fractional proof decomposition can automatically generate targeted unit tests for rare bugs, such as Anthropic's approximate top-K bug, without relying on prior bug reproducer code.
Vision Transformers (ViTs) can be enhanced with Post Hoc Registers (PH-Reg), a self-distillation method that integrates register tokens into existing models without full retraining, effectively mitigating the impact of artifact tokens on performance.
Conditional memory via Engram introduces a new axis of sparsity for large language models, enabling O(1) lookup and optimizing the balance between neural computation and static memory, as detailed in the arXiv paper.
Physical AI merges foundation models with robotics, showcasing advancements in Vision-Language-Action (VLA) models and world models, reflecting a rapid evolution in the field over the past 18 months.
Parallel Context-of-Experts Decoding (Pced) offers a novel solution to the retrieval-augmented generation dilemma by enabling cross-document reasoning without the need for shared attention mechanisms, thus enhancing efficiency in processing multiple documents.
Models often fail in production due to their reliance on correlations rather than understanding causal mechanisms, leading to poor decision-making despite high accuracy metrics.
Semantic caching for LLMs is complex, requiring a dual-layer architecture that combines exact hash matching with vector similarity search for effective performance.
Multiplex Thinking introduces a stochastic soft reasoning mechanism that samples K candidate tokens, aggregating their embeddings into a single multiplex token, enhancing reasoning efficiency without increasing sequence length.
Agentic Context Evolution (ACE) introduces a dynamic framework that optimally balances retrieval and reasoning, enhancing performance in knowledge-intensive tasks by reducing irrelevant context noise.
VideoHEDGE introduces a novel framework for detecting hallucinations in video question answering, leveraging entropy-based reliability estimation to enhance accuracy in Video-VLMs.