ML Times

High-Fidelity Simultaneous Speech-to-Speech Translation

Hibiki is a decoder-only model that enables simultaneous speech translation by processing source and target speech in real-time, producing both text and audio tokens for effective speech-to-speech translation.

ChatGPT creates phisher's paradise by serving the wrong URLs for major companies

AI chatbots, like ChatGPT, incorrectly provide URLs for major companies only 66% of the time, creating a fertile ground for phishing attacks. This vulnerability arises as criminals exploit the AI's tendency to generate misleading links, potentially leading users to fraudulent sites.

AV1@Scale: Film Grain Synthesis, The Awakening

Netflix's AV1 Film Grain Synthesis (FGS) enhances streaming quality by preserving the artistic integrity of film grain while achieving a 66% bitrate reduction, allowing for high-quality video delivery with less data usage.

About AI Evals

RAG is not dead; it remains vital for AI applications, but developers must distinguish between effective retrieval strategies and misleading marketing claims that oversimplify its utility. Understanding the nuances of Retrieval-Augmented Generation (RAG) is essential for making informed architectural decisions in AI systems.

AI for Scientific Search

AI4Research presents a comprehensive survey on the application of AI in scientific research, addressing the lack of unified perspectives and systematic classifications in this rapidly evolving field.

Ubuntu 25.10 Raises RISC-V Profile Requirements

Ubuntu 25.10 elevates its RISC-V profile requirements from RVA20 to RVA23, mandating essential features like Vector and Hypervisor extensions, which most current RISC-V devices lack, thus limiting compatibility with existing hardware.

The End of Moore's Law for AI? Gemini Flash Offers a Warning

Google's price increase for the Gemini 2.5 Flash model marks a pivotal shift in AI economics, indicating that the era of consistently decreasing costs for AI intelligence may be over, as operational costs are now dictated by hardware limitations and demand dynamics.

An Algorithm for a Better Bookshelf

A new algorithm for the bookshelf problem achieves an expected cost of log n × (log(log n))² per insertion, significantly improving upon the previous log1.5 n cost and approaching the theoretical lower limit of log n.

Paper with code is completely down

Paper with Code is currently completely down after experiencing prior spam issues, indicating a significant disruption in access to machine learning research resources.

Fast and Simplex: 2-Simplicial Attention in Triton

The 2-simplicial Transformer architecture enhances token efficiency by utilizing trilinear functions, outperforming standard Transformers in tasks like mathematics, coding, and reasoning.

Self-Correction Bench: Revealing and Addressing the Self-Correction Blind Spot in LLMs

Self-Correction Blind Spot: LLMs struggle to correct their own outputs, with a 64.5% average blind spot rate across 14 models, indicating a significant gap in their self-correction capabilities.

Ring Quantization: Achieving 90% on CIFAR-10 with 2-bit Networks

The Ring Quantization method achieves 90% accuracy on CIFAR-10 with 2-bit networks, demonstrating unexpected robustness at low bit-widths, particularly with deeper architectures like ResNet-32.

MOTIF: Modular Thinking via Reinforcement Fine-tuning in LLMs

MOTIF introduces a novel reinforcement learning method that enables large language models (LLMs) to generate reasoning tokens across multiple rounds, effectively overcoming the context size limitation inherent in traditional LLMs.

FlowSpec: Continuous Pipelined Speculative Decoding for Efficient Distributed LLM Inference

FlowSpec introduces a pipeline-parallel tree-based speculative decoding framework that enhances distributed LLM inference efficiency by prioritizing important tokens and managing drafts effectively.

MemAgent: Reshaping Long-Context LLM with Multi-Conv RL-based Memory Agent

MemAgent introduces a novel agent workflow that optimizes long-text processing by reading in segments and employing an overwrite strategy for memory updates, addressing the challenge of handling infinitely long documents efficiently.