ML Times
Mar 11, 2026
Yann LeCun's startup AMI has secured over $1 billion to develop AI world models that prioritize understanding the physical world over language, challenging the prevailing belief that scaling large language models (LLMs) will yield human-level intelligence.
David Noel Ng achieved the top position on the HuggingFace Open LLM Leaderboard by duplicating layers in a 72-billion parameter model without altering weights, demonstrating that layer duplication can enhance reasoning capabilities. This innovative approach, termed LLM Neuroanatomy, suggests that the internal structure of AI models can be manipulated for improved performance without traditional fine-tuning methods.
An autonomous agent exploited a SQL injection vulnerability in McKinsey's AI platform, Lilli, gaining full access to sensitive data within two hours without any insider knowledge or credentials.
Crawl entire websites with a single API call using the new
/crawlendpoint in Cloudflare's Browser Rendering, enabling automatic discovery and rendering of pages in multiple formats like HTML and JSON.Intel's Heracles chip accelerates fully homomorphic encryption (FHE) computing by 5,000 times, enhancing the efficiency of operations involving encrypted data.
bitnet.cpp is a cutting-edge inference framework for 1-bit LLMs, achieving 1.37x to 6.17x speedups on various CPU architectures while significantly reducing energy consumption by up to 82.2%.
RCLI is a local voice AI for macOS, enabling 43 actions via voice commands with sub-200ms latency and no reliance on cloud services, powered by the proprietary MetalRT GPU engine for optimal performance on Apple Silicon.
Wiz officially joins Google, enhancing its mission to secure cloud environments at the speed of AI, ensuring organizations can protect their innovations without sacrificing speed.
Duplicating a specific block of 7 middle layers in Qwen2-72B, without altering weights, led to a #1 ranking on the Open LLM Leaderboard, demonstrating the importance of preserving functional circuits in model architecture.
Agent Browser Protocol (ABP) transforms web browsing into a step machine, allowing agents to interact with a stable, frozen state of the web, enhancing automation efficiency.
TADA leverages text-acoustic synchronization to achieve fast and reliable speech generation, enhancing the efficiency of voice synthesis technologies.
A paper submitted to ICML is entirely AI-generated, raising concerns about adherence to guidelines that prohibit LLM assistance in writing or reviewing submissions.
Shadow APIs used in research can lead to performance divergence up to 47% and unpredictable safety behavior, raising concerns about the validity of findings in 187 academic papers that relied on these services.
NVIDIA and Thinking Machines Lab have forged a multiyear partnership to deploy gigawatt-scale NVIDIA Vera Rubin systems, enhancing frontier model training and customizable AI platforms.
AutoKernel autonomously optimizes GPU kernels for any PyTorch model by profiling, extracting, and iteratively refining bottleneck kernels using Triton, achieving significant performance improvements overnight.