Gemini 3.0 spotted in the wild through A/B testing
Gemini 3.0 has been observed in A/B testing via Google AI Studio, showcasing impressive SVG generation capabilities, particularly with an Xbox 360 controller image that outperformed existing models.
Claude Skills are awesome, maybe a bigger deal than MCP
Claude Skills enhance model performance by allowing Claude to load specific instructions and resources relevant to tasks, improving efficiency in specialized applications like Excel and brand compliance.
State of AI Report 2025
The State of AI Report 2025 reveals that OpenAI maintains a slight edge in AI development, while China's DeepSeek and others are rapidly closing the gap in reasoning and coding tasks, marking a significant shift in global AI leadership.
Bringing AI to the next generation of fusion energy
Google DeepMind is collaborating with Commonwealth Fusion Systems (CFS) to harness AI for advancing fusion energy, aiming to achieve net fusion energy with the SPARC tokamak, a significant step towards sustainable energy.
Nvidia DGX Spark and Apple Mac Studio = 4x Faster LLM Inference with EXO 1.0
EXO 1.0 achieves 4x faster LLM inference by leveraging the NVIDIA DGX Spark™ for prefill and Apple Mac Studio for decode, optimizing each phase's strengths.
[R] Plain English outperforms JSON for LLM tool calling: +18pp accuracy, -70% variance
Natural Language Tools (NLT) enhances LLM tool-call accuracy by +18 percentage points while reducing variance by 70% and token overhead by 31%, demonstrating a significant shift from structured JSON to natural language frameworks.
Flight Simulator for the Brain Reveals How We Learn and Why Minds Go Off Course
A new computer model, CogLinks, simulates brain decision-making and adaptability, revealing how misjudgments in context can lead to psychiatric disorders like schizophrenia and OCD.
LaSeR: Reinforcement Learning with Last-Token Self-Rewarding
LaSeR introduces a novel approach to Reinforcement Learning by simplifying self-verification into a last-token self-rewarding score, enhancing the efficiency of Large Language Models (LLMs) during reasoning tasks.
Attention Is All You Need for KV Cache in Diffusion LLMs
Elastic-Cache optimizes key-value (KV) cache management in diffusion large language models (DLMs) by adaptively refreshing caches based on token importance, significantly enhancing prediction accuracy and reducing decoding latency.
Show HN: Searchable compression for JSON – ~99% page skip and sub-ms lookups
SEE (Semantic Entropy Encoding) achieves a combined size of ≈19.5% of raw data while enabling searchability and low I/O, making it advantageous for workloads requiring quick access to compressed JSON data.
[R] Tensor Logic: The Language of AI
Tensor Logic (TL) aims to unify Deep Learning and Symbolic AI, offering a framework that integrates various AI models, including neural networks and graphical models, into a cohesive system.
Asking AI to build scrapers should be easy right?
Skyvern's AI can now autonomously write and maintain code, achieving a remarkable 2.7x cost reduction and 2.3x speed increase in automation tasks.
How a Gemma model helped discover a new potential cancer therapy pathway
Google's C2S-Scale 27B model, with 27 billion parameters, has successfully identified a novel cancer therapy pathway by predicting the synergistic effect of silmitasertib and low-dose interferon on antigen presentation.
[R][D] A Quiet Bias in DL’s Building Blocks with Big Consequences
Deep learning's building blocks, such as activation functions and optimisers, introduce a foundational bias that shapes network representation and reasoning, suggesting a new symmetry-based design axis for improved model performance.