Lambda Stack - CUDNN 8 Upgrade Query - Technical Help - DeepTalk - Deep Learning Community
Lambda Stack - CUDNN 8 Upgrade Query
post by subhrajeitbhowmick on Nov 6, 2020
Hi,
I use a Tensorbook and need to leverage TensorFlow GPU support for CUDA 11. Though the latest Lambda Stack upgrade switched my previous CUDA 10.2 to 11.1, the CUDNN version still remains 7.6.
Do we know of a timeline by when we can expect Lambda Stack to upgrade its CUDNN 7.6 to CUDNN 8.x?
Alternatively, is there a suggestion on how to upgrade it manually without breaking the Lambda Stack - ensuring no compatibility issues in future upgrades of the Lambda Stack?
Many thanks in advance!
Regards.
post by sabalaba on Nov 13, 2020
Lambda Stack with CuDNN 8 is coming shortly.
In the meantime, you should be able to link to CUDNN by modifying your LD_LIBRARY_PATH to include a path to the libcudnn.so.8 files.
post by subhrajeitbhowmick on Nov 14, 2020
Many thanks @sabalaba for your kind response!
I downloaded cuDNN 8 for Ubuntu 20.04 and added the LD_LIBRARY_PATH as per your suggestion.
However, while testing for TensorFlow’s GPU association, I received an error stating, “Could not load dynamic library ‘libcusolver.so.10’; dlerror: libcusolver.so.10: cannot open shared object file: No such file or directory”.
This might be resolved by completely removing all installed CUDA files and a fresh install of CUDA, but I fear this would lead to an incongruity in the existing Lambda Stack.
Note: I even tried with TensorFlow 2.4.0-rc1 whose pip packages are now built with CUDA11 and cuDNN 8.0.2.
Would you kindly be able to suggest a solution here?
post by perreiradasilva-m on Nov 18, 2020
Hello to all,
Just to add more details @sabalaba (because I faced the same problem), it seems there is however a problem with the current Lambda Stack version. On a fresh install, any PyTorch call to code that uses cuDNN throws an error:
"Could not load library libcudnn_cnn_train.so.8. Error: libcudnn_ops_train.so.8: cannot open shared object file: No such file or directory"
Please make sure libcudnn_cnn_train.so.8 is in your library path!
The only way to get things working is to set LD_LIBRARY_PATH accordingly:
export LD_LIBRARY_PATH=/usr/lib/python3/dist-packages/torch/lib/
But it seems not to be that lambda stack way.
I understand that cuDNN 8 is not yet supported but in that case why does a fresh install try to use cuDNN 8 by default?
post by sabalaba on Dec 1, 2020
Is there a reason that you’re using a pip installed version of PyTorch instead of the default Lambda Stack PyTorch?
The default Lambda Stack PyTorch has cuDNN built in and doesn’t throw that error.
>>> import torch
>>> torch.__path__
['/usr/lib/python3/dist-packages/torch']
>>> torch.__version__
'1.6.0'
post by perreiradasilva-m on Dec 2, 2020
Hello,
No, not using any pip installed version of PyTorch. Just Ubuntu 20.04 and a fresh install of Lambda Stack (nothing else added apart from JupyterHub). In that configuration I must export the torch location in LD_LIBRARY_PATH. Otherwise, I’ve got the reported error about libcudnn_cnn_train.so.8 not found :-(.
PyTorch itself works fine, just the cuDNN related part that fails. Reported torch version is, however, 1.7.0.
>>> import torch
>>> torch.__path__
['/usr/lib/python3/dist-packages/torch']
>>> torch.__version__
'1.7.0'
post by willkaes on Dec 6, 2020
I solved this problem by running this in Python:
import torch
torch.__path__
>>> ['/some/path/to/torch']
Then in terminal:
export LD_LIBRARY_PATH=$LD_LIBRARY_PATH:/some/path/to/torch
post by suvrat on Apr 23, 2021
Any updates on when it will be supported by default?