VRAM-hungry LSTM monster - Technical Help - DeepTalk - Deep Learning Community

VRAM-hungry LSTM monster

post by Movsisyan on Jun 8, 2022

Hello, I met an issue that seems to only be present in the lambda workstation so far.

Here is the link to the issue

post by markd on Jun 10, 2022

You have to love the name… VRAM-hungry LSTM monster.

I took a quick look. I was able to repeat the hang/dead kernel in Jupyter notebook and it failing on the command line.

From the command line I noticed it was not finding ‘libcudnn_ops_train.so.8’.

More information below the workaround.

And I was able to fix this by…

  1. Stopped jupyter notebook
  2. Set the LD_LIBRARY_PATH (which I thought was set in ld.so.conf.d
   $ export LD_LIBRARY_PATH=/usr/lib/python3/dist-packages/tensorflow:$LD_LIBRARY_PATH
  1. Start Jupyter notebook
   $ jupyter notebook
  1. Run the code as is in Jupyter Notebook
   $ web gui tool - which you obviously know 

From the command line I just commented out the two nvidia-smi, since I was watching both in nvtop and with:

$ nvidia-smi --query-gpu=index,pci.bus_id,fan.speed,utilization.gpu,utilization.memory,temperature.gpu,power.draw --format=csv -l

Here is the longer form of the message from the command line:

2022-06-09 23:35:37.010639: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1525] Created device /job:localhost/replica:0/task:0/device:GPU:1 with 8063 MB memory: → device: 1, name: NVIDIA GeForce RTX 3080, pci bus id: 0000:21:00.0, compute capability: 8.6
2022-06-09 23:35:38.611401: I tensorflow/stream_executor/cuda/cuda_dnn.cc:368] Loaded cuDNN version 8303

Could not load library libcudnn_adv_train.so.8. Error: libcudnn_ops_train.so.8: cannot open shared object file: No such file or directory

[lambda-dual:705714] *** Process received signal ***

[lambda-dual:705714] Signal: Aborted (6)

[lambda-dual:705714] Signal code: (-6)

… then it hung … until eventually the kernel in jupyter notebook died.

NOTE you can turn off much of the noisy tensorflow default messages with:

```bash
$ TF_CPP_MIN_LOG_LEVEL=3 python LSTM-hell.py

The workaround:

$ export LD_LIBRARY_PATH=/usr/lib/python3/dist-packages/tensorflow:$LD_LIBRARY_PATH
$ python LSTM-hell.py

… tensorflow noisy messages …

2022-06-09 23:37:27.934136: I tensorflow/core/common_runtime/gpu/gpu_device.cc:1525] Created device /job:localhost/replica:0/task:0/device:GPU:0 with 4138 MB memory: → device: 0, name: NVIDIA GeForce RTX 3080, pci bus id: 0000:01:00.0, compute capability: 8.6
2022-06-09 23:37:27.934442: I tensorflow/stream_executor/cuda/cuda_gpu_executor.cc:936] successful NUMA node read from SysFS had negative value (-1), but there must be at least one NUMA node, so returning NUMA node zero
...

post by Movsisyan on Jun 10, 2022

wsadmin@AIML1001:/usr/lib/python3/dist-packages$ find tensorflow*/ libcudnn_ops_train.so.8
find: ‘libcudnn_ops_train.so.8’: No such file or directory

Hello, thanks for the detailed response. What is your LD_LIBRARY_PATH? Mine was not set before running export LD_LIBRARY_PATH=/usr/lib/python3/dist-packages/tensorflow:$LD_LIBRARY_PATH so it became /usr/lib/python3/dist-packages/tensorflow:

And here is me running the workaround:

wsadmin@AIML1001:/home/mher/projects/Untitled Folder$ printenv | grep LD_
LD_LIBRARY_PATH=/usr/lib/python3/dist-packages/tensorflow:

post by markd on Jun 10, 2022

That is a good point.

I should have mentioned it is located in /usr/lib/python3/dist-packages/tensorflow if you have installed Lambda Stack for libcudnn_ops_train.so.8.

If you do not have lambda stack you would need to install cudnn from NVIDIA, where you need to register to get to the cudnn download page.

Cudnn: CUDA Deep Neural Network (cuDNN)  NVIDIA Developer

And you need to make sure wherever you install it the LD_LIBRARY_PATH can be seen. Also so that the version is compatible with the driver and CUDA you have installed.

If you are running inside of anaconda. Anaconda does not setup the library paths for packages.

You should be able to find the path, the likely locations would be:

LD_LIBRARY_PATH=${CONDA_PREFIX}/lib:${LD_LIBRARY_PATH} jupyter notebook

Let me know if that helps.

Or if you run directly from python (with the nvidia-smi lines commented out)

python LSTM-hell.py | tee -a LSTM-hell.txt

And send the LSTM-hell.txt

post by Movsisyan on Jun 10, 2022

Apologies if I’m messing up anywhere along the steps to the workaround but now I notice that I actually have the files under that directory:

wsadmin@AIML1001:/usr/lib/python3/dist-packages/tensorflow$ ls
_api      distribute   libcudnn_adv_infer.so.8  libcudnn_cnn_train.so.8  libcudnn.so.8                 __pycache__
compiler  include      libcudnn_adv_train.so.8  libcudnn_ops_infer.so.8  libtensorflow_framework.so.2  python
core      __init__.py  libcudnn_cnn_infer.so.8  libcudnn_ops_train.so.8  lite                          tools

And I updated the Lambda Stack a week ago (this issue persisted despite the upgrade) using the command mentioned in the official website:

sudo apt-get update && sudo apt-get dist-upgrade

post by markd on Jun 10, 2022

I think it will be easier working with you directly. Please send a email to support@lambdal.com - and I can work directly with you.

It is good you now have the libcudnn.so and the other parts of Lambda Stack.

And that is working correctly on my machine as long as I set the LD_LIBRARY_PATH.

There should not be any sudo required. The ‘tee -a’ was just to capture stdout so we can skip that.

export LD_LIBRARY_PATH=/usr/lib/python3/dist-packages/tensorflow
LD_LIBRARY_PATH=/usr/lib/python3/dist-packages/tensorflow python LSTM-hell.py

post by Movsisyan on Jun 13, 2022

This worked! Thank you so much.