# Sep 20, 2025

## Articles

- **Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens**  
  Chain-of-Thought (CoT) reasoning in LLMs may be a superficial phenomenon, as it often relies on learned patterns from training data rather than genuine inferential processes, suggesting a need for critical evaluation of its effectiveness.

- **Hidden risk in Notion 3.0 AI agents: Web search tool abuse for data exfiltration**  
  Notion 3.0 AI Agents introduce a significant vulnerability through the web search tool, enabling attackers to exfiltrate sensitive data by crafting malicious queries that exploit the tool's input schema.

- **The LLM Lobotomy?**  
  LLM performance degradation is observed over time, with consistent testing revealing that models like gpt-4o-mini yield increasingly inaccurate responses despite unchanged inputs.

- **Micro-LEDs boost random number generation**  
  Micro-LEDs developed by KAUST researchers achieve an ultra-high random number generation rate of 9.375 Gbit/s, utilizing intensity fluctuations in their spontaneous emission.

- **Philips announces digital pathology scanner with native DICOM JPEG XL output**  
  Philips has launched the Pathology Scanner SGi, the first to offer native DICOM JPEG XL output, which reduces file sizes by up to 50% while maintaining high image quality, enhancing data management in pathology labs.

- **Supporting Our AI Overlords: Redesigning Data Systems to Be Agent-First**  
  Large Language Model (LLM) agents are set to dominate data systems, necessitating a shift towards agent-first architectures that can efficiently handle their unique workloads, termed *agentic speculation*.

- **LLM-Deflate: Extracting LLMs into Datasets**  
  LLM-Deflate enables the extraction of structured datasets from trained large language models (LLMs), effectively reversing the lossy compression of knowledge into reusable training data, with promising results from three open-source models.

- **Evals in 2025: benchmarks to build models people can use**  
  In 2025, evaluations will focus on building models that are not just intelligent but also practical, emphasizing their utility in real-world applications. This shift is driven by the need for models that effectively manage ambiguity, follow instructions, and adapt to dynamic environments, as highlighted by recent reports from [Anthropic](https://www.anthropic.com/research/anthropic-economic-index-september-2025-report) and [OpenAI](https://cdn.openai.com/pdf/a253471f-8260-40c6-a2cc-aa93fe9f142e/economic-research-chatgpt-usage-paper.pdf).

- **Building sub-100ms autocompletion for JetBrains IDEs**  
  Next-edit autocomplete for JetBrains IDEs achieves sub-100ms response times by leveraging Diff-based Syntax-Aware FIM to ensure suggestions are contextually relevant and syntactically valid, enhancing developer trust and efficiency.

- **Overcoming accuracy limitations of Analog In-Memory Computing hardware**  
  Analog in-memory computing (AIMC) enhances neural network inference speed and power efficiency but faces challenges like noisy computations and input/output quantization constraints, limiting conventional LLM performance on AIMC hardware.

- **MiniGrid DoorKeys Benchmark Active Inference**  
  The Active Inference Framework demonstrates impressive performance on the MiniGrid DoorKeys (MG-DK) benchmark, achieving an average of <19 steps for an 8x8 grid and <60 steps for a 16x16 grid without extensive training or benchmarking.

- **Benchmarked EpilepsyBench #1 winner - found 27x performance gap, now training Bi-Mamba-2 fix**  
  The SeizureTransformer achieved a remarkable 26.89 FA/24h, revealing a 27x performance gap compared to previous benchmarks on the Temple EEG dataset, showcasing significant advancements in EEG machine learning.

- **Governed multi-expert aka (GME)**  
  The Governed Multi-Expert (GME) architecture transforms a single large language model into a dynamic team of specialists using Low-Rank Adaptation (LoRA) modules, enhancing response quality and safety while optimizing computational resources.

- **TorchAO Quantized Models and Quantization Recipes Now Available on HuggingFace Hub**  
  PyTorch has launched native quantized models like [Phi4-mini-instruct](https://huggingface.co/collections/pytorch/torchao-quantized-phi-4-mini-instruct-681566f123acc6fed345cb1a) and [Qwen3](https://huggingface.co/collections/pytorch/torchao-quantized-qwen3-6823b58cf63390161de2643b) that utilize int4 and float8 quantization for efficient inference on various devices, achieving minimal quality loss compared to bfloat16 models.
