Meta and Mark Zuckerberg face a lawsuit from five publishers and author Scott Turow, alleging they illegally copied millions of works to train their AI systems, specifically the Llama model, which is described as one of the largest copyright infringements in history.
Accelerating Gemma 4: faster inference with multi-token prediction drafters
Multi-Token Prediction (MTP) drafters enhance Gemma 4 models, achieving up to 3x faster inference without compromising output quality, thanks to a specialized speculative decoding architecture.
Computer Use Is 45x More Expensive Than Structured APIs
Computer use is 45 times more expensive than structured APIs, as evidenced by a benchmark showing 53 steps and 551k tokens for a vision agent versus only 8 calls and 12k tokens for an API agent.
Production AI very different from the demos
Production AI reveals a stark contrast to demos, as token usage surged unexpectedly due to longer, more complex customer queries, necessitating context retrieval that doubled input length on each call.
Why SSMs struggle in parameter-constrained training: empirical findings at 25M parameters
SSMs exhibit a 3.26x compression disadvantage compared to attention QKV, severely impacting their performance in parameter-constrained environments like the Parameter Golf competition.
Simple Meta-Harness on Islo.dev
A meta-harness enhances LLM agents by utilizing a proposer agent that analyzes up to 10M tokens of raw execution traces to improve the harness automatically, addressing the critical issue of diagnostic context.
Is there a notable increase in demand for privacy-preserving AI/ML with the advent of LLMs?
The demand for privacy-preserving AI/ML has surged alongside the rise of LLMs, driven by concerns over user de-anonymization as highlighted in this paper.
TritonSigmoid: A fast, padding-aware sigmoid attention kernel for GPUs
TritonSigmoid is an open-sourced, padding-aware sigmoid attention kernel designed for GPUs, enhancing performance in single-cell foundation models by allowing simultaneous attention to multiple genes without wasting compute on empty positions.
Visual graph classification for blockchain security: Experiences fine-tuning Qwen2-VL on AMD MI300X
Visual graph classification using Qwen2-VL-2B-Instruct effectively identifies malicious transaction patterns in blockchain security, leveraging a Vision-Language Model (VLM) to recognize topological signatures in 2D graph layouts.
Perceptual Flow Network for Visually Grounded Reasoning
PFlowNet introduces a novel approach to visual reasoning by decoupling perception from reasoning, enhancing interpretability and effectiveness beyond traditional geometric priors.
I'm Scared About Biological Computing
The emergence of biological computing—where human neurons are trained to perform tasks like playing DOOM—raises profound ethical questions about consciousness and the implications of creating a living computer.
QLoRA Fine-Tuning of Qwen2.5-1.5B for CEFR English Proficiency Classification (A1–C2)
QLoRA fine-tuning of Qwen2.5-1.5B achieved an impressive accuracy of 84.9% in classifying English texts across six CEFR levels (A1–C2), utilizing only ~0.28% of model parameters for training.
NVIDIA and ServiceNow Partner on New Autonomous AI Agents for Enterprises
NVIDIA and ServiceNow are launching specialized autonomous AI agents through Project Arc, designed to enhance enterprise workflows with governance and security, enabling complex task execution beyond traditional automation.