The $6 R1 competitor demonstrates that a small model can achieve near state-of-the-art performance with minimal data, highlighting a significant breakthrough in AI efficiency.
Ingesting PDFs and why Gemini 2.0 changes everything
Gemini 2.0 revolutionizes the processing of millions of PDFs, enhancing data extraction and analysis capabilities significantly.
Gemini 2.0 is now available to everyone
Gemini 2.0 introduces three new models: Flash, Flash-Lite, and Pro Experimental, enhancing performance and accessibility for developers and users alike.
DeepRAG: Thinking to retrieval step by step for large language models
DeepRAG enhances retrieval-augmented reasoning by modeling it as a Markov Decision Process (MDP), allowing for strategic and adaptive retrieval that reduces factual hallucinations in Large Language Models (LLMs).
How to scale your model: A systems view of LLMs on TPUs
Scaling LLMs on TPUs requires understanding hardware interactions and optimizing parallelism to achieve efficient training and inference, addressing common questions about costs and memory needs.
Catgrad: A categorical deep learning compiler
catgrad is a deep learning framework that leverages category theory to statically compile models, enabling training loops to operate independently of any deep learning framework, including catgrad itself.
Prediction Games
The Netflix Prize revolutionized machine learning by incentivizing improvements to its recommendation system, attracting over 5,000 teams and demonstrating that simple algorithms could yield significant results, as evidenced by a submission that utilized stochastic gradient descent for matrix factorization.
[R] On the Reasoning Capacity of AI Models and How to Quantify It
Recent advances in Large Language Models (LLMs) reveal that while they excel in benchmarks like GPQA and MMLU, they struggle with complex reasoning tasks, necessitating new evaluation methodologies to assess their true capabilities.
[R] Transformer-Squared: Self-adaptive LLMs
Transformer-Squared enables LLMs to dynamically adjust weights during inference, enhancing adaptability for unseen tasks through a two-pass mechanism that utilizes task-specific 'expert' vectors trained via reinforcement learning.
[D] How to Scale Your Model: A Systems View of LLMs on TPUs
Scaling LLMs is grounded in understanding system resources—compute, memory, and bandwidth—allowing for precise calculations of cost, runtime, and optimal parallelism strategies.
[P] Open-source library to generate ML models using natural language
smolmodels is an open-source Python library that generates ML models from natural language descriptions, utilizing graph search and LLM code generation to optimize model training for specific tasks. GitHub Repository
OCR Crypto Stealers in Google Play and App Store
SparkCat malware has infiltrated both Google Play and the App Store, utilizing an OCR model to extract sensitive crypto wallet recovery phrases from users' image galleries, marking a significant escalation in mobile malware threats.
Gemini 2.0 is now available to everyone
Gemini 2.0 introduces Flash, Flash-Lite, and Pro Experimental models, enhancing performance and accessibility for developers and users alike.
NVIDIA Blackwell Now Generally Available in the Cloud
CoreWeave has launched the first NVIDIA GB200 NVL72 instances, enabling high-performance cloud computing for AI reasoning, which requires significant compute power and optimized software for real-time results.
How GeForce RTX 50 Series GPUs Are Built to Supercharge Generative AI on PCs
NVIDIA's GeForce RTX 50 Series GPUs, leveraging the Blackwell architecture, deliver up to 3,352 AI trillion operations per second (TOPS), significantly enhancing local AI performance for developers and enthusiasts.