# Sep 13, 2025

- **Larry Ellison emphasizes that inference, not training, is the key revenue driver** in AI, highlighting a potential shift in focus towards efficient and scalable deployment of models.

- **Demis Hassabis**, CEO of Google's DeepMind, asserts that **"learning how to learn"** will be essential for future generations to adapt to the rapid changes brought by **AI** in education and the workplace.

- **MCP-Agent** enables the creation of **scalable deep research agents** by integrating state-of-the-art LLMs with MCP servers, allowing for complex task execution and context management through tool calls.

- The **Generalized Windowed Operation (GWO)** framework unifies deep learning operations by decomposing them into three components: **Path**, **Shape**, and **Weight**, enhancing operational locality and feature importance.

- **Qwen3's rapid expansion** includes the launch of 32 open-source models optimized for Apple’s MLX framework, enhancing AI performance on devices like MacBooks and iPhones through efficient quantization techniques.

- OpenAI's proposed solution to **AI hallucinations**—requiring models to assess their confidence before responding—could lead to a significant decline in user satisfaction, as it may prompt models to say "I don't know" for up to **30%** of queries, fundamentally altering user experience.

- The new paper on **Long Horizon Execution** reveals that the perception of **AI progress slowing down** is an _illusion_, with test-time scaling yielding significant advantages for long horizon autonomous agents. [Read the paper here](https://www.alphaxiv.org/abs/2509.09677).

- **K2-Think** claimed a **state-of-the-art (SOTA)** small model, but subsequent analysis raised significant doubts about its validity, as detailed in a [debunking post](https://www.sri.inf.ethz.ch/blog/k2think).

- **Real-time monitoring** of AI agents using Timeplus can effectively detect **hallucinations** during tasks, such as illegal moves in chess, ensuring errors are caught immediately to prevent serious consequences in critical applications like finance and healthcare.

- **PyTorch 2.8** introduces **native XCCL support** for Intel GPUs, enabling seamless distributed training and enhancing the user experience by integrating the Intel® oneAPI Collective Communications Library directly into the framework.

- **PyTorch and vLLM's integration enhances generative AI applications** by implementing Prefill/Decode Disaggregation, which optimizes inference efficiency, reducing latency and increasing throughput for large-scale systems like Meta's LLM products.
