# Mar 21, 2025

- **SoftBank Group** will acquire **Ampere Computing** for **$6.5 billion**, enhancing its AI infrastructure and semiconductor capabilities, crucial for advancing **Artificial Super Intelligence**.

- **SmolDocling** is an ultra-compact vision-language model designed for **end-to-end document conversion**, utilizing a novel markup format called **DocTags** to capture all page elements in context and location, with only **256M parameters**.

- **Gemma 3** is touted as the **most powerful AI model** operable on a single GPU, capable of interpreting **text, images, and short videos** across over 35 languages, enhancing its utility for developers.

- **Jagged Flash Attention** optimizes large-scale recommendation systems by achieving **up to 9× speedup** and **22× memory reduction** compared to traditional dense attention methods, enhancing both performance and scalability.

- **E-graphs and the egglog library** enable efficient term rewriting and optimization of Python expressions, allowing for the compilation of these expressions into MLIR, which can significantly enhance performance in numerical computations. The source code is available on [GitHub](https://github.com/sdiehl/mlir-egglog).

- **Particle accelerators** may revolutionize chip manufacturing by enabling the etching of **nano-scale designs** with high-speed electrons, potentially surpassing current optical lithography methods.

- **RNNs** excel in solving NC1 complexity class problems, yet the field has largely shifted to Transformers, which lack this capability, despite years of investment in optimizing hardware for them.

- **FlashTokenizer** is the **fastest tokenizer library** for LLM inference, developed in C++ to optimize speed and accuracy, significantly enhancing performance in natural language processing tasks.

- **Semi-supervised learning (SSL)** can be enhanced by using **parameter-efficient fine-tuning (PEFT)** with labeled data, often matching SSL performance without unlabeled data, highlighting a shift in approach for leveraging **vision foundation models (VFMs)**.

- **Piccolo** introduces a novel **end-to-end graph processing accelerator** that utilizes **fine-grained in-memory random scatter-gather** to significantly reduce off-chip traffic, enhancing efficiency beyond traditional methods.

- The paper reveals that **small sliding windows yield flat centroids**, while **large windows** effectively form **interval clusters**, providing a mathematical basis for observed clustering behaviors in time series data.

- **TULIP** enhances vision-language models by integrating **contrastive learning** with **masked feature prediction**, effectively addressing the "seeing half a scene" problem prevalent in models like CLIP.

- **Yandex Research** has successfully distilled **SD3.5 Large/Medium** into fast few-step generators, achieving performance comparable to two-step sampling while outperforming other distillation methods within the same compute budget.

- The **Open Power AI Consortium**, led by EPRI and NVIDIA, aims to develop **open AI models** to enhance electricity generation and distribution, addressing challenges like distributed energy resources and grid load growth.

- **GTC 2025** will showcase cutting-edge advancements in **AI, robotics, and accelerated computing**, featuring influential speakers like **Yann LeCun** and **Frances Arnold**, who will challenge conventional thinking and inspire innovation.
