ML Times
Apr 6, 2026
Claude Code's recent updates have severely degraded its ability to handle complex engineering tasks, with a significant regression in performance noted since February, as evidenced by a drop in effective thinking depth and increased error rates in user interactions.
Google Gemma 4's 26B-A4B model leverages a mixture-of-experts architecture, activating only 4B parameters per forward pass, making it efficient for local inference on standard hardware like a MacBook Pro with 48 GB of memory. This design allows for high performance without the need for extensive computational resources, achieving 51 tokens per second during operation.
Age verification laws in Brazil, the UK, and the US are driving the creation of biometric identity verification systems that function as mass surveillance tools, with Peter Thiel's investments linking surveillance analytics and identity verification companies, thereby establishing a coordinated legislative demand across borders.
Quantum-resistant cryptography is now deemed urgent, with experts suggesting that 2029 is a critical deadline for migration, as recent research indicates that breaking 256-bit elliptic curves may soon be feasible with fewer qubits than previously thought.
nanocodeis a library designed for training Claude Code using Constitutional AI, enabling users to create their own models with a focus on agentic coding behaviors. This project leverages JAX and is optimized for TPU training, allowing for efficient model development at a low cost, approximately $200 for a 1.3B parameter model in about 9 hours.Dante-2B is a 2.1B parameter bilingual LLM trained from scratch on Italian and English, achieving coherent Italian text generation in just 16 days using 2× H200 GPUs without relying on existing models.
Deep Extract revolutionizes structured extraction by employing an agent-in-the-loop system that autonomously verifies and corrects its output, achieving 99-100% accuracy on complex documents.
Koru kernels outperform idiomatic C, Rust, and Zig, achieving performance within 1% of hand-specialized C while simplifying the programming process by embedding optimization-relevant semantics directly into the code structure.
Fused MoE dispatch kernel in pure Triton outperforms Stanford's Megablocks at inference batch sizes, achieving 131% and 124% speed improvements at 32 and 128 tokens, respectively, while maintaining compatibility across multiple models.
The method is primarily novel, achieving superior performance against all baselines, including those it was not expected to surpass, which has garnered surprise and recognition in the field.
Freestyle offers a robust platform for managing AI-generated code, enabling rapid VM provisioning in under 700ms and live forking of VMs without downtime, enhancing development efficiency.
Gradient-boosted attention enhances standard attention mechanisms by introducing a second pass that corrects prediction errors, akin to Friedman's gradient boosting machine.
The Hallucination-as-Cue Framework reveals that reinforcement learning (RL) can enhance Multimodal Large Language Models (MLLMs) by leveraging model hallucination, challenging traditional views on visual reasoning capabilities.
This survey highlights augmentation strategies for large language models (LLMs), focusing on the structured context provided during inference, including methods like in-context learning, Retrieval-Augmented Generation (RAG), and CausalRAG.
JoyAI-LLM Flash redefines the trade-off between performance and token efficiency in sub-50B parameter models, utilizing a novel RL algorithm called FiberPO for enhanced stability in policy optimization.