# Feb 6, 2025

## Ingesting PDFs and why Gemini 2.0 changes everything
- **Gemini 2.0** revolutionizes the processing of **millions of PDFs**, enhancing data extraction and analysis capabilities significantly.

## S1: A $6 R1 competitor?
- The **$6 R1 competitor** demonstrates that a small model can achieve near state-of-the-art performance with minimal data, highlighting a significant breakthrough in AI efficiency.

## Gemini 2.0 is now available to everyone
- **Gemini 2.0** introduces three new models: **Flash**, **Flash-Lite**, and **Pro Experimental**, enhancing performance and accessibility for developers and users alike.

## How Deepseek trained their R1 models, and how frontier LLMs are trained today.
- **Deepseek's innovative Mixture of Experts configuration** utilizes a high sparsity factor of **8/256**, significantly enhancing model performance by ensuring all experts contribute across tasks through an auxiliary loss mechanism.

## US Cloud soon illegal in EU? US punches first hole in EU-US Data Deal
- **Trump's removal of Democratic members from the PCLOB jeopardizes the EU-US Data Transfer Framework**, raising concerns about the independence of US oversight bodies and the adequacy of data protection for EU citizens.

## Pre-Trained Large Language Models Use Fourier Features for Addition (2024)
- **Pre-trained LLMs utilize Fourier features** to compute addition, leveraging both low-frequency and high-frequency dimensions in their hidden states for effective arithmetic reasoning.

## R1 Computer Use
- **r1-computer-use** leverages **large-scale Reinforcement Learning** to enhance computer interaction, utilizing a **neural reward model** to assess the correctness of actions taken by the agent in various environments like file systems and web browsers.

## Transformer-Squared: Self-adaptive LLMs
- **Transformer-Squared** enables LLMs to **dynamically adjust weights during inference**, enhancing adaptability for unseen tasks through a two-pass mechanism that utilizes task-specific 'expert' vectors trained via reinforcement learning.

## Harmonic Loss Trains Interpretable AI Models
- **Harmonic loss** offers a novel approach by utilizing **Euclidean distance** instead of the traditional inner product, leading to improved model performance and interpretability.

## Evaluating Code Embeddings
- **Vector-based code retrieval** is essential for modern coding assistants, yet evaluating the quality of embedding models remains challenging due to a lack of diverse, high-quality benchmarking datasets and methodologies for their creation.

## How to Scale Your Model: A Systems View of LLMs on TPUs
- **Scaling LLMs** is grounded in understanding **system resources**—compute, memory, and bandwidth—allowing for precise calculations of cost, runtime, and optimal parallelism strategies.

## Consistency Models: Why doesn’t the model collapse?
- **Consistency models** maintain performance by balancing **consistency distillation** and **consistency training losses**, preventing collapse into trivial outputs like all zeros.

## DeepRAG: A Markov Decision Process Framework for Step-by-Step Retrieval-Augmented Reasoning
- **DeepRAG** revolutionizes retrieval-augmented generation by employing a **"Think-before-Retrieval"** architecture, which enhances reasoning accuracy through a structured, step-by-step approach before information retrieval.

## Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning
- **Hybrid representation** of reasoning using latent discrete tokens from VQ-VAE reduces input length and computational demands, enhancing efficiency in Large Language Models (LLMs).
