The Lambda Deep Learning Blog | engineering

The Lambda Deep Learning Blog

Recent Posts

Kimi K2 Thinking: what 200+ tool calls mean for production

Published on February 4, 2026
by Lea Alcantara

TL;DR: Kimi K2 Thinking is Moonshot AI's open-source reasoning model, scoring 44.9% on Humanity's Last Exam with the ability to chain 200-300 sequential tool calls.


How to deploy ML jobs on Lambda Cloud with SkyPilot

Published on January 20, 2026
by Cody Brownstein

TL;DR: SkyPilot is an open-source orchestration tool that automates ML job deployment on Lambda Cloud. This tutorial covers installation, configuration, and usage.


2025 AI wrapped

Published on January 13, 2026
by Lea Alcantara

2025 was a year of momentum in AI. Intelligence progressed through new, innovative methods. Open-source communities released competitive models. Research labs continued to push boundaries.


JAX on NVIDIA GPUs Part 2: A practical guide for ML engineers

Published on January 8, 2026
by Jessica Nicholson

This guide demonstrates how to scale JAX-based LLM training from a single GPU to multi-node clusters on NVIDIA Blackwell infrastructure. We present a detailed walkthrough for effective model training.


How to serve Kimi-K2-Instruct on Lambda with vLLM

Published on December 22, 2025
by Zach Mueller

When your model doesn’t fit on a single GPU, you suddenly need to target multiple GPUs on a single machine, configure a serving stack that actually uses all resources effectively.


JAX on NVIDIA GPUs Part 1: Fundamentals and best practices

Published on September 19, 2025
by Jessica Nicholson

JAX unlocks distinct advantages on GPUs: automatic kernel fusion via XLA, composable transformations, and hardware-agnostic code that moves between different computational setups seamlessly.


LLM performance up 15.4%: MLPerf v5.1 confirms NVIDIA HGX B200 on Lambda is built for enterprise inference

Published on September 9, 2025
by Anket Sah

Inference at scale is still too slow. Large models often stall under real-world load, burning time, compute, and user trust. That’s the problem we set out to address with unique optimization strategies.


The Essential Guide to GPUs for AI, Training and Inference

Published on August 20, 2025
by Jessica Nicholson

Introduction: Graphics Processing Units (GPUs) were originally designed to handle computer graphics, like making video games look realistic or helping services like Netflix stream smoothly. This article dives into their use in AI.


Introducing Next-Gen Training and Inference at Scale on Lambda Instances with NVIDIA Blackwell

Published on August 12, 2025
by Anket Sah

NVIDIA Blackwell GPUs are now available as 8x Lambda Instances On-Demand, featuring the powerful NVIDIA HGX™ B200 in addition to our trusted lineup.


Beginners Guide to Reasoning in AI

Published on August 7, 2025
by Jessica Nicholson

If you've been anywhere near LLMs lately, you've probably heard the word "reasoning" thrown around more than a frisbee at a college campus. GPT-4 can "reason" introduces its functionalities and applications in AI.


Introducing NVIDIA SHARP on Lambda 1CC: Next-Gen Performance for Distributed AI Workloads

Published on July 29, 2025
by Anket Sah

Lambda’s 1-Click Clusters (1CC) provide AI teams with streamlined access to scalable, multi-node GPU clusters, simplifying the complexity of distributed AI workloads.


Accelerate Your AI Workflow with FP4 Quantization on Lambda

Published on July 16, 2025
by Anket Sah

As AI models grow in complexity and size, the demand for efficient computation becomes paramount. FP4 (4-bit Floating Point) precision emerges as a feasible solution for scaling.


Apriel 5B: ServiceNow’s Enterprise AI Trained and Deployed on Lambda

Published on June 20, 2025
by Anket Sah

In AI, scaling doesn’t always mean "bigger." We emphasize lean, efficient LLM design that maximizes performance while minimizing compute costs and resource use.


Partner Spotlight: Orchestrating large-scale agent training on Lambda with dstack and RAGEN

Published on June 5, 2025
by dstack

Lambda + dstack: Empowering your ML team with robust infrastructure for distributed reasoning agent training.


DeepSeek-R1-0528: The Open-Source Titan Now Live on Lambda’s Inference API

Published on June 4, 2025
by Anket Sah

DeepSeek has just leveled up. The latest release, DeepSeek-R1-0528, is now available on Lambda’s Inference API, delivering a formidable blend of mathematical intuitiveness and efficiency.