Replies - Activity - markd - DeepTalk - Deep Learning Community
Tensorbook not booting up
Changing /etc/X11 would not affect grub or linux, just the Desktop showing up.
To get to the BIOS you would select the ‘F2’ key for most machines, and the 2020 version of the tensorbook should be the ‘F2’ key.
Then you could be from a USB device also if you need to recover, once you are in the BIOS.
Webcam is not working
What system is this for? Is this a USB camera or is this on a notebook and which notebook? Commonly it is either:
- unplug/replug-in a USB Camera (so it reloads)
- Or confirm the driver is loaded
Check for drivers:
$ lsmod | grep video
Reload the driver would depending on camera/connection.
How to install CUDA 11?
Thank you! It is important to always list where you are running and your requirements. You have not listed which GPUs you are running on, etc. It is fairly simple just following the basic instructions Cody listed. The Drive is NOT CUDA. The Driver shows the maximum CUDA level it supports.
Compression by jpeg2000 (nvJPEG2000 library)
Yes, you just download and install it. nvJPEG2000 downloads There are examples on: GitHub - NVIDIA/CUDALibrarySamples
- Select the nvJPEG2000
- It is a few years old, but looks like it was updated 3 weeks ago
Unable to receive display from USB C or Display Ports on Vector Quad
There is not a USB-C on these GPUs, so you are likely just connecting to motherboard USB-C used for normal USB data. Do you see the BIOS on the screen while it boots? Or does it only go blank after the OS gets started?
General Issues:
- Cable Issues (try a known working cable)
- Wrong setting
A6000 (Ampere) temperatures
These are available on Linux and on some motherboards that have IPMI (WRX80E). Temperatures in the chassis will vary greatly depending on:
- The amount of memory
- CPU type
- Fans and cooling (which normally depends on the overall hardware)
- If you added hardware
- How many GPUs
- Fan speed adjustments
- Airflow
Best way to switch OS from Windows to Linux
It is pretty straight forward on Ubuntu. For workstations, go to Ubuntu Install Instructions this takes you step-by-step through downloading, creating a USB image and installing Ubuntu. Then install Lambda Stack:
wget -nv -O- https://lambdalabs.com/install-lambda-stack
"ChildFailedError" fine tuning Video-Llama and Video-ChatGPT
The simple answer is you are running distributed, and the parent process is telling you that one of the child processes failed. It is not clear for which reason, but it could be:
- Insufficient resources for the child process (GPU, GPU memory, CPU, memory)
- Perhaps if this is a remote host it could be...
Lambda Quad (Ubuntu) no display signal at all
You can go to Lambda Support and open a ticket. It looks like you have done that and they are helping already. To save time/money the first step is:
- Clear the CMOS - just to make sure it is not a simple fix.
Then it would likely require shipping, debugging. A Qcode of 00 is likely the motherboard.
Own PC but GPU from Lambda with SSH?
Yes it is possible. Installing stable diffusion on the cloud is the easier way though, instructions are available and many have used this. Automatic 1111 implementation is a bit more finicky scripts; another option is a docker image on the cloud. The important thing to consider is the speed of your deployment.
Python 3.8 and tensorflow 2.10
Was this meant to be part of a thread? And is this your own machine you are running a VM on? The cloud is using VMs. So yes you will need Python and TensorFlow installed for TensorFlow-based project. You can ‘version’ using Docker, Python venv or Anaconda/miniconda. And that can be done on both setups:...
How to upgrade to Ubuntu 20.04 while keeping Lambda Stack?
That is from a few years ago but works for removal, and the script should be used for the install. How do I remove and reinstall Lambda Stack? Is the instructions. It will work for Ubuntu 20.04 LTS or Ubuntu 22.04 LTS (either desktop or server). There could be other packages that...
Frequent segmentation faults
SEGV or segfaults are caused when a code is attempting to access invalid or protected memory. The reasons for this can be:
Software:
- Application has a bug and trying to allocate/use an incorrect memory address
- This can be due to data the code uses, corrupt library, bug in the code itself
Hardware:
- Issues with RAM could potentially lead to this...
Best practices for other package management without breaking my Lambda Stack?
Yes, it is best practice not to use pip in your main account, but to use versioning (as Cody mentioned).
pip -v list | egrep -v “/usr/lib/python3/dist-packages”
- This will show all packages that are in your current environment not from Lambda.
- It is best to use: Docker, Python venv, Anaconda/Miniconda.
Ubuntu won't boot to desktop after disconnecting and then re-connecting SSD
It would be hard to say without the journalctl -xb or the logs. You should be able to hit ‘return’ to get to the prompt. (This looks like single user boot or ‘recovery mode’). It could be complaining for many reasons, but it looks like the motherboard, CPU, memory all seem happy. It could be...
Unable to determine the device handle for GPU0000:21:00.0: Unknown Error
running ‘sudo nvidia-bug-report.sh’ would produce a ‘nvidia-bug-report.log.gz’ which would be sufficient to see what is failing. Then normally swap the two GPUs; this is to confirm:
- GPUs are reseated
- That the failure follows GPU versus the PCI slot
Boot hangs with the root file system on /dev/sda2 requires a manual fsck
Be careful of this. It is fine to do to get up and running. But it is concerning. If you need to do a fsck to get it back:
- Do a backup of that partition
- Install ‘smartmontools’:
$ sudo apt install smartmontools
- Run smartctl to check the health of that drive; it may be failing and perhaps should be replaced...
Can Anaconda coexist with Python installed via Lambda Stack?
Yes, Anaconda is isolated for python/pip/tensorflow/pytorch/cudnn/cuda. So with Anaconda or Python venv you would need to properly install CUDA with the PyTorch (see www.pytorch.org getting started matrix) or TensorFlow. Also, if you need cuDNN make sure you set the path for the version you installed.
Tensorflow and torch not found
A common issue is if other things are installed in ~/.local or /usr/local. Doing the following would likely provide the information of what is seen:
$ pip -v list | egrep “torch|tensorflow”
So if you are inside a virtual environment (Anaconda or python venv) it would ‘hide’ the system-installed packages...
Lambda Stack version archive
The latest Lambda Stack comes with 525.116.04, which supports up to CUDA 12.0. The current CUDA in use is 11.8. The Driver version is commonly newer than the CUDA version, as most issues are in CUDA versus the driver and the driver supports a few major versions back. You can check the CUDA version...
Where is cudnn.h please?
Just a update. At least on Ubuntu 22.04 it looks like there is a ‘nvidia-cudnn’ which does include the C header files.
$ apt search nvidia-cudnn
$ sudo apt install nvidia-cudnn
Model runs on A10, but not H100
I hacked this a little bit, testing the normal TensorFlow design issues and workarounds. I ended up:
$ TF_FORCE_GPU_ALLOW_GROWTH='true' python test.py
And added the following to the code:
from tensorflow.compat.v1 import ConfigProto
from tensorflow.compat.v1 import InteractiveSession
config = ConfigProto()
config.gpu_options.allow_growth = True
session = InteractiveSession(config=config)
Ubuntu 22.04 support timeline
I would not use Ubuntu’s --allow-third-party, as that will attempt to install their versions of the NVIDIA drivers, CUDA, etc. Versus the version that PyTorch and TensorFlow are built with. There would be more information in the logs about the conflicts. Also you should do:
$ sudo apt update
Black screen on TensorFlow
This likely means that the GUI terminal session is in use. Hit one of the combinations ‘Control-Alt-F2’ to ‘Control-Alt-F5’ to go to a command line session. Login, clean up and reconfirm install. One reason is some commands create a xorg.conf...
Lambda Blade 8 RTX GPU, stuck at AF code
If it is a blade it should have IPMI. Ideally you know how to use IPMI and have the network setup. The chassis would have one ‘Ethernet’ labeled IPMI LAN or similar; it should get an address via dhcp. And the chassis should have a password. Where can I find my server's IPMI (BMC) password?
Local storage to keep data
There is persistent storage is available only at the Texas location. Unfortunately that location is often full.
Include JAX in LambdaStack
We are looking into adding JAX. It requires more testing.
Recovery Image for Razer x Lambda Tensorbook
There was a updated image, just use: Recovery Image Documentation It will list the current images there.
NVIDIA driver not loaded after GCC upgrade
I would recommend sending an email to ‘support@lambdal.com’ with the ‘nvidia-bug-report.log.gz’ from ‘sudo nvidia-bug-report.sh’ and the ‘dpkg.txt’ from ‘sudo dpkg --list > dpkg.txt’.
When updating kernel the NVIDIA driver disappears
If you update the kernel (or the driver) the driver kernel module needs to be rebuilt against that specific kernel. We have tested/confirmed the kernel-nvidia driver on the latest Ubuntu 22.04 LTS desktop...
Change workstation cables
For power cables, it will depend on the specific setup your new Building uses. For example, this is the one we recommend in the US for 2000W power supplies....
Ubuntu 22.04 network error
You can see if a IP address is assigned with:
$ ip address show
To test the network you can ping an IP address directly, for example (Google's DNS):
$ ping -c 3 8.8.8.8
Upgrade graphics cards
As Cody mentioned, the considerations you need to make are:
Power - it depends on how many GPUs and how much power you may need.
- 4x 2080’s is about 1000 Watts (250 Watts each)
- 4x 3090’s is about 1400 Watts (350 Watts each) + spikes above.
Cooling/Thermal - this is probably fine, but may not be optimal for performance...
GPU instance not booting
Please send a note to ‘support@lambdal.com’ with the email address you used for the instance, explaining the above. And the billing can be cleared up also.
Tensorflow: "No protocol specified"
Another way to find out what is happening is run your code with:
$ strace python -c 'import tensorflow as tf' >& tensorflow-strace.txt
You can send that output file here or contact ‘support@lambdal.com’ with that file and nvidia-bug-report.log.gz from ‘sudo nvidia-bug-report.sh’.
Low Disk Space on filesystem.root
The normal Ubuntu install has all the software in the same partition. (/, /usr, /tmp, /var, and /home). There is also a separate small /boot/efi partition that you do not need to worry about unless it gets full...
Pytorch and conda on Lambda Workstation RTX 3090
I would not take payment. I would feel that is a conflict of interest. As I work at Lambda, and eventually we may have more consulting, my first priority is just to get you up and running. Since it is not under a support contract lets do it ‘after hours’.