The Lambda Deep Learning Blog | engineering
Kimi K2 Thinking: what 200+ tool calls mean for production
TL;DR: Kimi K2 Thinking is Moonshot AI's open-source reasoning model, scoring 44.9% on Humanity's Last Exam with the ability to chain 200-300 sequential tool calls, greatly enhancing production capabilities.
How to deploy ML jobs on Lambda Cloud with SkyPilot
TL;DR: SkyPilot is an open-source orchestration tool that automates ML job deployment on Lambda Cloud. This tutorial covers installation, configuration, and best practices to seamlessly run ML workloads in the cloud.
2025 AI wrapped
2025 was a year of momentum in AI. Intelligence progressed through new, innovative methods. Open-source communities released competitive models. Research labs explored new frontiers, pushing the limits of AI capabilities and applications.
JAX on NVIDIA GPUs Part 2: A practical guide for ML engineers
This guide demonstrates how to scale JAX-based LLM training from a single GPU to multi-node clusters on NVIDIA Blackwell infrastructure. We present best practices and performance optimization techniques.
How to serve Kimi-K2-Instruct on Lambda with vLLM
When your model doesn’t fit on a single GPU, you suddenly need to target multiple GPUs on a single machine, configure a serving stack that actually uses all available resources efficiently.
JAX on NVIDIA GPUs Part 1: Fundamentals and best practices
JAX unlocks distinct advantages on GPUs: automatic kernel fusion via XLA, composable transformations, and hardware-agnostic code that moves seamlessly between platforms, making it a powerful choice for ML engineers.
LLM performance up 15.4%: MLPerf v5.1 confirms NVIDIA HGX B200 on Lambda is built for enterprise inference
Inference at scale is still too slow. Large models often stall under real-world load, burning time, compute, and user trust. We address this challenge through improved architecture and optimization.
The Essential Guide to GPUs for AI, Training and Inference
Introduction Graphics Processing Units (GPUs) were originally designed to handle computer graphics, like making video games look realistic or helping Netflix stream smoothly. Today, they are essential tools in AI model training and inference.
Introducing Next-Gen Training and Inference at Scale on Lambda Instances with NVIDIA Blackwell
NVIDIA Blackwell GPUs are now available as 8x Lambda Instances On-Demand, featuring the powerful NVIDIA HGX™ B200 in addition to our advanced lineup to optimize AI workflows at scale.
Beginners Guide to Reasoning in AI
If you've been anywhere near LLMs lately, you've probably heard the word "reasoning" thrown around more than a frisbee at a college campus. GPT-4 can "reason" through tasks effectively, showcasing the evolving capabilities of AI.
Introducing NVIDIA SHARP on Lambda 1CC: Next-Gen Performance for Distributed AI Workloads
Lambda’s 1-Click Clusters(1CC) provide AI teams with streamlined access to scalable, multi-node GPU clusters, cutting through the complexity of distributed AI workloads for better performance.
Accelerate Your AI Workflow with FP4 Quantization on Lambda
As AI models grow in complexity and size, the demand for efficient computation becomes paramount. FP4 (4-bit Floating Point) precision is emerging as an essential strategy to enhance AI model performance.
Apriel 5B: ServiceNow’s Enterprise AI Trained and Deployed on Lambda
In AI, scaling doesn’t always mean “bigger.” We focus on lean, efficient LLM design that maximizes performance while minimizing compute costs for enterprise applications.
Partner Spotlight: Orchestrating large-scale agent training on Lambda with dstack and RAGEN
Lambda + dstack: Empowering your ML team with robust infrastructure for distributed reasoning agent training, enhancing capabilities and efficiency in AI tasks.
DeepSeek-R1-0528: The Open-Source Titan Now Live on Lambda’s Inference API
DeepSeek has advanced its offering. The latest release, DeepSeek-R1-0528, is now available on Lambda’s Inference API, providing a powerful blend of computational capabilities.