All - Activity - yanos - DeepTalk - Deep Learning Community

CUDA drivers not working on fresh 1xH100 instance

This was the same issue. Since we have a low availability on the H100s, this kind of hardware failures show up more often because only the failed instances are the ones that eventually become available. We will hopefully have a fix for this soon so that they don’t become available after the first …


Too many open files

Hello! @pathos00011, @yapee23 Can you send us an example of the command you run that causes this to happen, so we can reproduce this? Best, Yanos


Tensorrt support?

Hello @nrkarp, I was able to install TensorRT in a Lambda cloud instance using the following commands.

$ pip install -U pip
$ pip install wheel
$ pip install tensorrt

$ python
Python 3.8.10 (default, Nov 14 2022, 12:59:47)
>>> import tensorrt
>>> print(tensorrt.__version__)
8.6.1

Let me know i…


CUDA drivers not working on fresh 1xH100 instance

@cudahell, @md23 Thank you very much for letting us know about this. This was a hardware issue on the GPUs. You definitely helped a lot of other users from having the same frustration. We are also working on adding a check to auto-blocklist GPUs with similar hardware failures.


CUDA drivers not working on fresh 1xH100 instance

Hello @cudahell, I was not able to reproduce this. If you upgrade the kernel you need to reboot the node because the kernel modules and the libraries have different versions but this should be a different error message:

$ nvidia-smi
Failed to initialize NVML: Driver/library version mismatch

You …


Cannot Run Falcon40B on H100

@Gadersd, please start a ticket in https://support.lambdalabs.com and we can look more into it.


Uploading/Downloading directory with scp

Hello Leo, Uploading or downloading directories with SCP using the -r flag should work fine. Can you send the command you run and the error output you get? If you prefer you can open a ticket here: https://support.lambdalabs.com Best, Yanos


Warnings on H100 instance running PyTorch

Hello @keithlostracco! This warning messages can be safely ignored. See also here. Best, Yanos


Cannot Run Falcon40B on H100

Hi @Gadersd! Can you send the error you are getting on the H100 instance? Best, Yanos


How to terminate instance from shell?

From the instance’s terminal, you can only shut down the operating system but you cannot terminate the instance so that it stops being charged. One workaround is the following and you can run it from any PC or even from inside the instance.

Generate a Lambda Cloud API Key GPU Cloud Login | Lamb…


Where is cudnn.h please?

The Lambda Stack has the CUDNN library for Tensorflow and Pytorch software in the following directories:

/usr/lib/python3/dist-packages/torch/lib/libcudnn.so.8
/usr/lib/python3/dist-packages/tensorflow/libcudnn.so.8

There is no “cudnn.h” file by default with Lambda Stack if that is what you are lo…


Tranformer Engine installation fails

Hi! The easiest way to have Nvidia’s Transformer Engine library is the NGC Pytorch Docker image. The Transformer Engine library is preinstalled in the PyTorch container in versions 22.09 and later on NVIDIA GPU Cloud. (ref. Installation — Transformer Engine 0.9.0 documentation) First, you will …


Best Practices(?): common tactics for server setups & data storage

Hello @JimVincent and @comodoro! Here is a simple way to deploy your software stack on the Lambda Cloud using Ansible. eg. create “ready_to_work.yml”

# you can run this as follows:
# ansible-playbook --ssh-common-args='-o StrictHostKeyChecking=no' ready_to_work.yml -i lambda_cloud_instance_ip…

Open MPI warning - no preset parameters were found

Hi Den, In your case, you can definitely ignore this. The Mellanox card doesn’t take any part in your workflow. Best, Yanos


Open MPI warning - no preset parameters were found

Hello Den, That error means that the Mellanox card isn’t known by Open MPI lib. As the message suggests this may be affecting performance. If you’re not doing multi-node training this shouldn’t be a problem for Open MPI not to be able to optimize for your ethernet NIC. Can you tell me more about…


How to use GPU's when training a model?

Can you see the GPU(s) when you run “nvidia-smi”? Also, do you by any chance have overridden any of the Lambda Stack installed Python packages? Can you share the output of “pip -v list”?