Lambda MFU.pdf

Reproducible MFU optimization techniques to boost your training efficiency

Lambda’s optimizations consistently drive >25% efficiency gains for foundation model training on NVIDIA HGX B200 and GB300 NVL72.

Executive summary

In this paper, we demonstrate that co-optimized hardware configuration and training stack can substantially increase Model FLOPS Utilization (MFU) on NVIDIA HGX B200 and NVIDIA GB300 NVL72 platforms. Across 10 representative workloads, we observe >25% MFU uplift, with a median improvement of approximately 35% over the industry baseline.

Results span 8-128 NVIDIA HGX B200 and NVIDIA GB300 NVL72 GPUs, 8B-405B parameter models, and 8k-32k sequence lengths, and are achieved without changes to model architecture or datasets. The efficiency gains were driven by Lambda’s ML experts:

Together, these results demonstrate that system-level optimization can unlock a substantial fraction of otherwise underutilized accelerator performance, even for mature model architectures.

The efficiency challenge: Bridging the gap between peak and actual utilization

Training foundation models on large clusters spans many layers of software and hardware, all working together, and each layer has opportunities for optimization. Even small performance improvements lead to better hardware utilization, faster training times, and faster evolution of the state-of-the-art models.

To measure the extent of optimizations, the ML community at large uses Model FLOPS Utilization (MFU) as the metric to optimize training workloads. MFU is defined as how close a training run comes to theoretical maximum performance:

$$ M F U = O b s e r v e d F L O P S / P e a k T h e o r e t i c a l F L O P S $$

MFU is very easy to compute. Peak Theoretical FLOPS is a static value retrieved from open GPU datasheets, and the Observed FLOPS is computed using throughput:

$$ O b s e r v e d F L O P S = (F L O P / t o k e n) (t o k e n s / s e c o n d) $$

Notably, the FLOP/token is model-dependent value indicating how many floating point operations are involved in one training step for your model, or simply, “how big your model is.” Dense models typically use the following formula:

$$ F L O P / t o k e n = 6 \left(N - N _ {e m b}\right) + 1 2 L H Q S $$

MFU on NVIDIA HGX B200 and NVIDIA GB300 NVL72

Our goal is to measure improvements in MFU that are directly attributable to Lambda’s training stack and optimizations, starting from a clean baseline. In order to fully exercise the hardware and optimization processes, we chose three standard Llama 3.1 models and trained them across Lambda’s NVIDIA HGX B200 and NVIDIA GB300 NVL72 clusters:

For each model, we optimized multiple context lengths, as both short and long contexts are essential to training high-quality LLMs:

This setup allows us to optimize both memory and compute-bound workloads, as well as workloads that run across nodes and racks.

Training stack

All experiments use a consistent training stack to ensure comparability across runs:

Baseline (industry standard and TorchTitan)

To ensure fair, reproducible comparisons, all performance comparisons start from an identical, controlled baseline to isolate the effects of batch size, model size, and optimization strategies on performance.

Our baseline is both the industry standard 35-45% (baseline 40%, as median of the range) and the torchtitan@e7ee95a codebase with the following rubric:

Results

Across all evaluated configurations, our results consistently improve MFU relative to industry standard and TorchTitan baselines. Compared to industry standards, we achieve a median MFU improvement of 1.36x across all workloads. This improvement is consistent across sequence lengths, with MFU reaching up to 55% at 8k pretraining and increasing to up to 60% at 16k–32k.

TorchTitan baseline MFU(8k sequence length) TorchTitan with Lambda optimizations MFU(8k seq length) Lambda Uplift vs. TorchTitan baseline MFU
Llama-3.1-8b BF168x NVIDIA HGX B200 44.46% 55.34% 1.24x
Llama-3.1-70b BF1616x NVIDIA HGX B200 23.83% 50.20% 2.11x
Llama-3.1-405b BF16128xNVIDIA GB300 NVL72 28.62% 52.69% 1.84x

Scaling with sequence length

As sequence length increases, activation memory, attention cost, and communication volume also increase. Despite these pressures, optimized configurations maintain higher MFU across all evaluated sequence lengths:

Training step profiling

We took a profile from a single training step, calculated the time each kernel took to complete, and provided a breakdown of the main kernel calls.

Takeaways:

How to optimize your MFU

We applied a set of targeted optimizations designed to keep the GPU as busy as possible. These techniques address suboptimal software and hardware configurations that commonly limit MFU in large-scale training.

The gains in our results come from four main areas:

  1. Kernel fusion
  2. Memory optimizations
  3. Communication/computation overlap
  4. System-level configuration

Communication/computation overlap

Even with efficient kernels and memory layouts, MFU suffers when computation and communication serialize. The following optimizations focus on overlapping communication with useful compute wherever possible.

  1. Kernel fusion
    • Merging the QKV computation
    • Fused optimizer kernels
    • torch.compile for model and loss functions
  2. Memory optimizations
    • Activation checkpointing (AC)
    • Tensor parallelism (TP)
  3. Communication/computation overlap
    • FSDP prefetching
    • Pinned memory

What this means for your training stack

If your large-scale training runs consistently operate below 40% MFU on large models, you’re leaving both performance and budget on the table. Meaningful throughput improvements require keeping existing GPUs busy with useful matrix multiplications, rather than recomputation, communication, or bookkeeping.

Appendix: calculating MFU

MFU is defined as “the ratio of the observed throughput (tokens-per-second) relative to the theoretical maximum throughput of a system operating at peak FLOPS.”

Observed FLOPS FLOPS (Floating Point Operations per Second) is calculated using:

$$ O b s e r v e d F L O P S = (F L O P / t o k e n) (t o k e n s / s e c o n d) $$

Peak Theoretical FLOPS The NVIDIA Blackwell datasheet specifies that the NVIDIA HGX B200 system performs 4.5 PFLOPS for INT8 operations, and the NVIDIA GB300 NVL72 rack performs 5 PFLOPS for bfloat16 operations.

Example

To put these metrics in context, consider training a Llama 8B model on a single NVIDIA B200 GPU using BF16 precision with sequence length 8K...

In summary, optimizing your configuration and leveraging targeted techniques can significantly enhance the efficiency of your training workloads.