ML Times

Feb 6, 2025

Ingesting PDFs and why Gemini 2.0 changes everything

S1: A $6 R1 competitor?

Gemini 2.0 is now available to everyone

How Deepseek trained their R1 models, and how frontier LLMs are trained today.

US Cloud soon illegal in EU? US punches first hole in EU-US Data Deal

Pre-Trained Large Language Models Use Fourier Features for Addition (2024)

R1 Computer Use

Transformer-Squared: Self-adaptive LLMs

Harmonic Loss Trains Interpretable AI Models

Evaluating Code Embeddings

How to Scale Your Model: A Systems View of LLMs on TPUs

Consistency Models: Why doesn’t the model collapse?

DeepRAG: A Markov Decision Process Framework for Step-by-Step Retrieval-Augmented Reasoning

Token Assorted: Mixing Latent and Text Tokens for Improved Language Model Reasoning