The Lambda Deep Learning Blog | Chuan Li
The Lambda Deep Learning Blog
Building at the speed of research: Lambda at CVPR 2026
Published on June 9, 2026 by Chuan Li and Ksenia Anske
Every year, CVPR draws the researchers defining what AI can see, understand, and act on. This year in Denver, more than 9,000 attendees showed up with over ...
ICLR 2026: 12 papers on making AI systems reliable, efficient, and secure
Published on April 23, 2026 by Chuan Li
A 7B agent that beats GPT-4o. Lossless weight compression that speeds up inference by 177%. An arena where 23 teams battled across 103,000 adversarial rounds. ...
From bigger models to better intelligence: what NeurIPS 2025 tells us about progress
Published on December 15, 2025 by Chuan Li
NeurIPS has always been a mirror: it doesn’t just reflect what the community is building, it reveals what the community is starting to believe. In 2025, that ...
Benchmarking ZeRO-Inference on the NVIDIA GH200 Grace Hopper Superchip
Published on December 20, 2023 by Chuan Li
This blog explores the synergy of DeepSpeed’s ZeRO-Inference, a technology designed to make large AI model inference more accessible and cost-effective, with ...
.png)
Unleashing the power of Transformers with NVIDIA Transformer Engine
Published on November 21, 2023 by Chuan Li
In this blog, Lambda showcases the capabilities of NVIDIA’s Transformer Engine, a cutting-edge library that accelerates the performance of transformer models ...
DeepChat 3-Step Training At Scale: Lambda’s Instances of NVIDIA H100 SXM5 vs A100 SXM4
Published on October 12, 2023 by Chuan Li
GPU benchmarks on Lambda’s offering of the NVIDIA H100 SXM5 vs the NVIDIA A100 SXM4 using DeepChat’s 3-step training example.
.png)
How FlashAttention-2 Accelerates LLMs on NVIDIA H100 and A100 GPUs
Published on August 24, 2023 by Chuan Li
This blog post walks you through how to use FlashAttention-2 on Lambda Cloud and outlines NVIDIA H100 vs NVIDIA A100 benchmark results for training GPT-3-style ...
How To Use mpirun to Launch a LLaMA Inference Job Across Multiple Cloud Instances
Published on March 14, 2023 by Chuan Li
.png)
Hugging Face x Lambda: Whisper Fine-Tuning Event
Published on December 1, 2022 by Chuan Li
Lambda is thrilled to team up with Hugging Face, a community platform that enables users to build, train, and deploy ML models based on open source code, for a ...
NVIDIA GeForce RTX 4090 vs RTX 3090 Deep Learning Benchmark
Published on October 31, 2022 by Chuan Li
Available October 2022, the NVIDIA® GeForce RTX 4090 is the newest GPU for gamers, creators, students, and researchers. In this post, we benchmark RTX 4090 to ...
NVIDIA H100 Tensor Core GPU - Deep Learning Performance Analysis
Published on October 5, 2022 by Chuan Li
We have seen groundbreaking progress in machine learning over the last couple of years. At the same time, massive usage of GPU infrastructure has become key to ...
Multi node PyTorch Distributed Training Guide For People In A Hurry
Published on August 26, 2022 by Chuan Li
Setting Up A Kubernetes Run:AI Cluster on Lambda Cloud
Published on June 3, 2022 by Chuan Li
If you're interested in training the next large transformer like DALL-E, Imagen, or BERT, a single GPU (or even single 8x GPU instance!) might not be enough ...
Best GPU for Deep Learning in 2022 (so far)
Published on February 28, 2022 by Chuan Li
NVIDIA A40 Deep Learning Benchmarks
Published on November 30, 2021 by Chuan Li
NVIDIA® A40 GPUs are now available on Lambda Scalar servers. In this post, we benchmark the A40 with 48 GB of GDDR6 VRAM to assess its training performance ...