NVIDIA A40 Deep Learning Benchmarks

NVIDIA A40 Deep Learning Benchmarks

November 30, 2021• 4 min read

NVIDIA® A40 GPUs are now available on Lambda Scalar servers. In this post, we benchmark the A40 with 48 GB of GDDR6 VRAM to assess its training performance using PyTorch and TensorFlow. We then compare it against the NVIDIA V100, RTX 8000, RTX 6000, and RTX 5000.

NVIDIA A40* Highlights

PyTorch ConvNet training speed

PyTorch language model training speed

PyTorch "32-bit" multi-GPU training scalability

We also tested the scalability of the A40 for multi-GPU training. To minimize system bottlenecks, our test server was equipped with two AMD EPYC 7763 64-Core Processors, 512 GB of system memory, and an NVME SSD.

The charts below show that the A40 achieved near perfect scaling from one to eight GPUs.

PyTorch "16-bit" multi-GPU training scalability

PyTorch benchmark software stack

Note: The GPUs were tested using NVIDIA PyTorch containers. The NVIDIA V100, RTX 8000, RTX 6000, RTX 5000, and RTX 4000 were tested with pytorch:20.01-py3. All other GPUs were benchmarked using pytorch:20.10-py3. While the performance impact of testing with different container versions is likely minimal, for completeness we are working on re-testing a wider range of GPUs using the latest containers and software. Stay tuned for an update.

Lambda's PyTorch benchmark code is available at the GitHub repo here.