Lambda MFU.pdf
Reproducible MFU optimization techniques to boost your training efficiency
Lambda’s optimizations consistently drive >25% efficiency gains for foundation model training on NVIDIA HGX B200 and GB300 NVL72.
Executive summary
In this paper, we demonstrate that co-optimized hardware configuration and training stack can substantially increase Model FLOPS Utilization (MFU) on NVIDIA HGX B200 and NVIDIA GB300 NVL72 platforms. Across 10 representative workloads, we observe >25% MFU uplift, with a median improvement of approximately 35% over the industry baseline.
Results span 8-128 NVIDIA HGX B200 and NVIDIA GB300 NVL72 GPUs, 8B-405B parameter models, and 8k-32k sequence lengths, and are achieved without changes to model architecture or datasets. The efficiency gains were driven by Lambda’s ML experts:
- Utilizing open source training stacks and extrapolating configuration best practices from previous-generation GPUs.
- Engineering high-performance infrastructure:
- NVIDIA HGX B200—400 Gb/s NVIDIA Quantum-2 InfiniBand connectivity per GPU for cross-node communication, and fifth-generation NVIDIA NVLink providing 1.8 TB/s of GPU-to-GPU bandwidth for intra-node communication.
- NVIDIA GB300 NVL72 system—800 Gb/s per GPU NVIDIA Quantum-X800 InfiniBand for cross-rack communication, and 130 TB/s total NVLink bandwidth per GPU across the rack-wide NVLink fabric.
Together, these results demonstrate that system-level optimization can unlock a substantial fraction of otherwise underutilized accelerator performance, even for mature model architectures.
The efficiency challenge: Bridging the gap between peak and actual utilization
Training foundation models on large clusters spans many layers of software and hardware, all working together, and each layer has opportunities for optimization. Even small performance improvements lead to better hardware utilization, faster training times, and faster evolution of the state-of-the-art models.
To measure the extent of optimizations, the ML community at large uses Model FLOPS Utilization (MFU) as the metric to optimize training workloads. MFU is defined as how close a training run comes to theoretical maximum performance:
$$ M F U = O b s e r v e d F L O P S / P e a k T h e o r e t i c a l F L O P S $$
MFU is very easy to compute. Peak Theoretical FLOPS is a static value retrieved from open GPU datasheets, and the Observed FLOPS is computed using throughput:
$$ O b s e r v e d F L O P S = (F L O P / t o k e n) (t o k e n s / s e c o n d) $$
Notably, the FLOP/token is model-dependent value indicating how many floating point operations are involved in one training step for your model, or simply, “how big your model is.” Dense models typically use the following formula:
$$ F L O P / t o k e n = 6 \left(N - N _ {e m b}\right) + 1 2 L H Q S $$
MFU on NVIDIA HGX B200 and NVIDIA GB300 NVL72
Our goal is to measure improvements in MFU that are directly attributable to Lambda’s training stack and optimizations, starting from a clean baseline. In order to fully exercise the hardware and optimization processes, we chose three standard Llama 3.1 models and trained them across Lambda’s NVIDIA HGX B200 and NVIDIA GB300 NVL72 clusters:
- Llama 3.1 8B on 8x NVIDIA HGX B200
- Llama 3.1 70B on 16x NVIDIA HGX B200
- Llama 3.1 405B on 2x NVIDIA GB300 NVL72
For each model, we optimized multiple context lengths, as both short and long contexts are essential to training high-quality LLMs:
- 8k tokens (8,192)
- 16k tokens (16,384)
- 32k tokens (32,768)
This setup allows us to optimize both memory and compute-bound workloads, as well as workloads that run across nodes and racks.
Training stack
All experiments use a consistent training stack to ensure comparability across runs:
- Framework: TorchTitan
- Precision: BF16 and BF16 + FP8
- Dataset: AllenAI C4, about 300GB of common crawl web data, streamed from Hugging Face
Baseline (industry standard and TorchTitan)
To ensure fair, reproducible comparisons, all performance comparisons start from an identical, controlled baseline to isolate the effects of batch size, model size, and optimization strategies on performance.
Our baseline is both the industry standard 35-45% (baseline 40%, as median of the range) and the torchtitan@e7ee95a codebase with the following rubric:
- Out-of-the-box TorchTitan configs
- Global batch size increased only to saturate GPU RAM
- No v-boost
- NVIDIA cuDNN attention
Results
Across all evaluated configurations, our results consistently improve MFU relative to industry standard and TorchTitan baselines. Compared to industry standards, we achieve a median MFU improvement of 1.36x across all workloads. This improvement is consistent across sequence lengths, with MFU reaching up to 55% at 8k pretraining and increasing to up to 60% at 16k–32k.
| TorchTitan baseline MFU(8k sequence length) | TorchTitan with Lambda optimizations MFU(8k seq length) | Lambda Uplift vs. TorchTitan baseline MFU | |
|---|---|---|---|
| Llama-3.1-8b BF168x NVIDIA HGX B200 | 44.46% | 55.34% | 1.24x |
| Llama-3.1-70b BF1616x NVIDIA HGX B200 | 23.83% | 50.20% | 2.11x |
| Llama-3.1-405b BF16128xNVIDIA GB300 NVL72 | 28.62% | 52.69% | 1.84x |
Scaling with sequence length
As sequence length increases, activation memory, attention cost, and communication volume also increase. Despite these pressures, optimized configurations maintain higher MFU across all evaluated sequence lengths:
- At 16K and 32K tokens, MFU continues to scale upward rather than collapsing under memory pressure.
- At longer sequence lengths, MFU approaches ~60%, indicating effective overlap of computation, communication, and memory access.
Training step profiling
We took a profile from a single training step, calculated the time each kernel took to complete, and provided a breakdown of the main kernel calls.
Takeaways:
- Matrix multiplication increased from 56% to 68% of train step time. We attribute this to kernel fusion by torch.compile and the larger batch size.
- Elementwise addition went from 14% of time to 1.5%.
- Attention in the forward pass was optimized from 7% to 3%, while attention in the backward pass increased from 11.5% to 15.4%.
How to optimize your MFU
We applied a set of targeted optimizations designed to keep the GPU as busy as possible. These techniques address suboptimal software and hardware configurations that commonly limit MFU in large-scale training.
The gains in our results come from four main areas:
- Kernel fusion
- Memory optimizations
- Communication/computation overlap
- System-level configuration
Communication/computation overlap
Even with efficient kernels and memory layouts, MFU suffers when computation and communication serialize. The following optimizations focus on overlapping communication with useful compute wherever possible.
- Kernel fusion
- Merging the QKV computation
- Fused optimizer kernels
- torch.compile for model and loss functions
- Memory optimizations
- Activation checkpointing (AC)
- Tensor parallelism (TP)
- Communication/computation overlap
- FSDP prefetching
- Pinned memory
What this means for your training stack
If your large-scale training runs consistently operate below 40% MFU on large models, you’re leaving both performance and budget on the table. Meaningful throughput improvements require keeping existing GPUs busy with useful matrix multiplications, rather than recomputation, communication, or bookkeeping.
Appendix: calculating MFU
MFU is defined as “the ratio of the observed throughput (tokens-per-second) relative to the theoretical maximum throughput of a system operating at peak FLOPS.”
Observed FLOPS FLOPS (Floating Point Operations per Second) is calculated using:
$$ O b s e r v e d F L O P S = (F L O P / t o k e n) (t o k e n s / s e c o n d) $$
Peak Theoretical FLOPS The NVIDIA Blackwell datasheet specifies that the NVIDIA HGX B200 system performs 4.5 PFLOPS for INT8 operations, and the NVIDIA GB300 NVL72 rack performs 5 PFLOPS for bfloat16 operations.
Example
To put these metrics in context, consider training a Llama 8B model on a single NVIDIA B200 GPU using BF16 precision with sequence length 8K...
In summary, optimizing your configuration and leveraging targeted techniques can significantly enhance the efficiency of your training workloads.