ML Times
Sep 20, 2025
Articles
Is Chain-of-Thought Reasoning of LLMs a Mirage? A Data Distribution Lens
Chain-of-Thought (CoT) reasoning in LLMs may be a superficial phenomenon, as it often relies on learned patterns from training data rather than genuine inferential processes, suggesting a need for critical evaluation of its effectiveness.Hidden risk in Notion 3.0 AI agents: Web search tool abuse for data exfiltration
Notion 3.0 AI Agents introduce a significant vulnerability through the web search tool, enabling attackers to exfiltrate sensitive data by crafting malicious queries that exploit the tool's input schema.The LLM Lobotomy?
LLM performance degradation is observed over time, with consistent testing revealing that models like gpt-4o-mini yield increasingly inaccurate responses despite unchanged inputs.Micro-LEDs boost random number generation
Micro-LEDs developed by KAUST researchers achieve an ultra-high random number generation rate of 9.375 Gbit/s, utilizing intensity fluctuations in their spontaneous emission.Philips announces digital pathology scanner with native DICOM JPEG XL output
Philips has launched the Pathology Scanner SGi, the first to offer native DICOM JPEG XL output, which reduces file sizes by up to 50% while maintaining high image quality, enhancing data management in pathology labs.Supporting Our AI Overlords: Redesigning Data Systems to Be Agent-First
Large Language Model (LLM) agents are set to dominate data systems, necessitating a shift towards agent-first architectures that can efficiently handle their unique workloads, termed agentic speculation.LLM-Deflate: Extracting LLMs into Datasets
LLM-Deflate enables the extraction of structured datasets from trained large language models (LLMs), effectively reversing the lossy compression of knowledge into reusable training data, with promising results from three open-source models.Evals in 2025: benchmarks to build models people can use
In 2025, evaluations will focus on building models that are not just intelligent but also practical, emphasizing their utility in real-world applications. This shift is driven by the need for models that effectively manage ambiguity, follow instructions, and adapt to dynamic environments, as highlighted by recent reports from Anthropic and OpenAI.Building sub-100ms autocompletion for JetBrains IDEs
Next-edit autocomplete for JetBrains IDEs achieves sub-100ms response times by leveraging Diff-based Syntax-Aware FIM to ensure suggestions are contextually relevant and syntactically valid, enhancing developer trust and efficiency.Overcoming accuracy limitations of Analog In-Memory Computing hardware
Analog in-memory computing (AIMC) enhances neural network inference speed and power efficiency but faces challenges like noisy computations and input/output quantization constraints, limiting conventional LLM performance on AIMC hardware.MiniGrid DoorKeys Benchmark Active Inference
The Active Inference Framework demonstrates impressive performance on the MiniGrid DoorKeys (MG-DK) benchmark, achieving an average of <19 steps for an 8x8 grid and <60 steps for a 16x16 grid without extensive training or benchmarking.Benchmarked EpilepsyBench #1 winner - found 27x performance gap, now training Bi-Mamba-2 fix
The SeizureTransformer achieved a remarkable 26.89 FA/24h, revealing a 27x performance gap compared to previous benchmarks on the Temple EEG dataset, showcasing significant advancements in EEG machine learning.Governed multi-expert aka (GME)
The Governed Multi-Expert (GME) architecture transforms a single large language model into a dynamic team of specialists using Low-Rank Adaptation (LoRA) modules, enhancing response quality and safety while optimizing computational resources.TorchAO Quantized Models and Quantization Recipes Now Available on HuggingFace Hub
PyTorch has launched native quantized models like Phi4-mini-instruct and Qwen3 that utilize int4 and float8 quantization for efficient inference on various devices, achieving minimal quality loss compared to bfloat16 models.