NVLink seems to prevent PyTorch training loop from starting - Technical Help - DeepTalk - Deep Learning Community
NVLink seems to prevent PyTorch training loop from starting
post by mnazemi on Jun 10, 2022
I have a Lambda server with eight A6000 GPUs and NVLink.
When I utilize PyTorch’s distributed data parallel (DDP) to train my models with two GPUs, NVLink is used successfully based on the performance counters. However, as soon as I increase the number of GPUs to three or more (up to eight), the training loop gets stuck at the very beginning (or possibly the first backward pass).
If I set the environment variable NCCL_P2P_DISABLE=1, I can use as many GPUs as I like, but I obviously don’t get the benefits of NVLink.
If I set NCCL_DEBUG=INFO, I get the following output:
...NCCL output...
I was wondering if anyone has had a similar issue and knows how to resolve it.
Thanks!
post by markd on Jun 14, 2022
If it is getting ‘stuck’ or ‘hanging’ do you have AMD CPUs?
And do you have the kernel command line set to iommu=soft?
Another issue I have seen was cross hosts and there were processes already running on the other GPU. Sometimes it took a little while for the previous job cleanup. (nvidia-smi to see what is running on the GPUs).
I do not use the NCCL_P2P_DISABLE=1
I normally run the DDP, NCCL test with:
NCCL_DEBUG=INFO NCCL_ALGO=Ring NCCL_NET_GDR_LEVEL=4 python …
To see your nvlink topology:
$ nvidia-smi topo -m
And to see the nvlink status:
$ nvidia-smi nvlink -s
You can use ‘CUDA_VISIBLE_DEVICES=0,1’ to run on the first two GPUs or ‘CUDA_VISIBLE_DEVICES=1,2’ to run on the second and third devices (if you are trying to compare).
You can see that with:
$ cat /proc/cmdline
To fix it you can:
$ sudo sed -iE ‘/^GRUB_CMDLINE.*DEFAULT/ s/"$/ iommu=soft"/’ /etc/default/grub
$ sudo update-grub
$ sudo reboot
post by mnazemi on Jun 15, 2022
Thank you for your detailed explanation. I should have mentioned that I use AMD CPUs and I don’t see this issue when using two GPUs.
The problem was indeed resolved by setting iommu=soft.
Given your explanation, I found a similar solution that relied on setting pci=noats, but it didn’t work for me.
It appears that there is currently no way to use the hardware MMU to get better performance.
Ironically, I don’t get a noticeable speedup when setting iommu=soft and NCCL_DEBUG=INFO NCCL_ALGO=Ring NCCL_NET_GDR_LEVEL=4. Here is the output of nvidia-smi topo -m:
...nvidia-smi output...