M5 achieves over 4x peak GPU compute performance for AI compared to M4, featuring a 10-core GPU with a Neural Accelerator in each core, enhancing both AI and graphics capabilities significantly.
iPad Pro with M5 chip
The new iPad Pro features the M5 chip, delivering up to 3.5x faster AI performance than the M4 and 5.6x faster than the M1, enhancing productivity and creativity through advanced GPU and CPU capabilities.
Claude Haiku 4.5
Claude Haiku 4.5 offers near-frontier coding performance at one-third the cost and twice the speed of its predecessor, Claude Sonnet 4, making it a game-changer for real-time AI applications.
Intel Announces Inference-Optimized Xe3P Graphics Card with 160GB VRAM
Intel's "Crescent Island" GPU features 160GB of LPDDR5x memory and is designed for AI inference, emphasizing performance-per-Watt, but will not be available until H2 2026.
How AI hears accents: An audible visualization of accent clusters
AI models, like BoldVoice's finetuned HuBERT, effectively cluster over 200 accents using a vast dataset of 30 million recordings, revealing unexpected relationships between accents based on geographic and social factors rather than linguistic taxonomy.
Just Talk to It – The No-Bs Way of Agentic Engineering
Agentic engineering has advanced to the point where it can autonomously write nearly all code, streamlining workflows and reducing unnecessary complexity in development processes.
Oops It's a kernel stack use-after-free: Exploiting Nvidia's GPU Linux drivers
Two critical vulnerabilities in NVIDIA's Linux GPU drivers, CVE-2025-23280 and CVE-2025-23300, allow local unprivileged processes to exploit kernel memory management, confirmed through a proof of concept that achieves kernel read and write primitives.
Prefix sum: 20 GB/s (2.6x baseline)
Delta coding enhances data compression by storing differences between successive values, achieving a remarkable 19.8 GB/s throughput— 1.8x faster than naive implementations and 2.6x faster than FastPFoR, thanks to optimized compute restructuring.
Beyond the SQLite Single-Writer Limitation with Concurrent Writes
Turso's concurrent writes enhance SQLite's capabilities, achieving up to 4x the write throughput while eliminating the common SQLITE_BUSY error, thus enabling more efficient database operations in modern applications.
[P] Nanonets-OCR2: An Open-Source Image-to-Markdown Model with LaTeX, Tables, flowcharts, handwritten docs, checkboxes & More
Nanonets-OCR2 is a cutting-edge model suite that excels in converting images to markdown, featuring capabilities like LaTeX recognition, signature isolation, and multilingual support for diverse document types.
Recursive Language Models (RLMs)
Recursive Language Models (RLMs) enable language models to decompose and recursively interact with input contexts of unbounded length, significantly improving performance on long-context tasks while mitigating "context rot."
Dr.LLM: Dynamic Layer Routing in LLMs
Dr.LLM introduces a dynamic routing framework for Large Language Models (LLMs) that allows pretrained models to adaptively skip, execute, or repeat layers, enhancing efficiency without significant accuracy loss.
[R]: Create a family of pre-trained LLMs of intermediate sizes from a single student-teacher pair
Boomerang distillation allows for the creation of a family of pre-trained LLMs of varying sizes by distilling a large teacher model into a smaller student and then reintegrating teacher layers, optimizing both performance and resource efficiency.
Memory as Action: Autonomous Context Curation for Long-Horizon Agentic Tasks
Memory-as-Action reframes working memory management in Large Language Models as a learnable capability, allowing agents to actively curate memory through explicit editing operations integrated into their policy.
[R] Verbalized Sampling: How to Mitigate Mode Collapse and Unlock LLM Diversity
Verbalized Sampling mitigates mode collapse in LLMs by prompting for probability distributions rather than single outputs, enhancing creative task diversity by 2.1x without sacrificing quality.