# What drivers/patches are necessary for the razer tensorbook?

I am not a LambdaLab person. I noticed that your github link above is not valid as it has space and hence it is starting to land at [b.com/eureka](http://b.com/eureka)… Good luck.

# Lambda Stack seems to be wiped out after Ubuntu update

Problem is solved. Since via

```bash
sudo apt-get remove lambda-stack-cuda && sudo apt-get install lambda-stack-cuda
```

This failed in the middle due to interruption and I have to do

```bash
sudo dpkg --configure -a
```

and reissue the above install of lambda-stack-cuda. This went through, I rebooted the system and...

# Lambda Stack CUDA and Nvidia drivers

Yes I also want to upgrade Cuda to 12.1 and compatible torch version. As per this Nvidia [Official Drivers | NVIDIA](https://www.nvidia.com/Download/index.aspx)
We can upgrade the driver via the above link and the torch via the [pytorch.org](http://pytorch.org/) page.
Is there a lambda-stack way to upgrade them?

# Hi Cody,

When I rebooted again, the system came back. I am able to connect.
How to check whether the

```bash
sudo apt-get install lambda-stack-cuda
```

is complete? When I do

```bash
sudo apt-cache policy lambda-stack-cuda
```

I get...

Thanks for reply.
On my LambdaBlade with Ubuntu 20.04 and 4X A6000, I wanted to upgrade the cuda stack.
So I issued the this command:

```bash
$ sudo apt-get remove lambda-stack-cuda && sudo apt-get install lambda-stack-cuda
```

After it successfully fetched 115 packages, extracted templates, read database....

# Just want to clarify.

System LambdaBlade with Ubuntu 20.04: Is there a difference between the command here [Getting-started Removing and Installing Lambda-stack-in-ubuntu](https://docs.lambdalabs.com/software/lambda-stack-and-recovery-images)?

Uninstall (purge) the existing Lambda Stack by running:

```bash
sudo rm -f /etc/apt/sources.list.d/{graphics,nvidia,cuda}* && \
dpkg ...
```

# Keyboard and Mouse not recognized when connected from a KVM switch

Both keyboard and mouse connected via KVM switchs are not being recognized by some of the recent Lambda machines. Two instances: 4 port KVM switch with HDMI + USB connections. It happened when I connected 3 Quads and 1 Vector via a 4 port KVM switch. The Vector machine doesn’t respond.

# Lambda Quad - How to Unbond NVLINK from P2P/SLI mode

Any update? I have a blade system with 4 A6000. What does the below output mean?

```bash
$ nvidia-smi nvlink -gt d -i 0,1,2,3
```

# Unable to determine the device handle for GPU, GPU is lost. Reboot the system to recover this GPU

As per the tech support, reboot did the trick. Hope this is not a persistent vibration related issue.

# I have a lambda blade server with 8 x RTX 3090s.

Today I lost one of the GPUs. Here is the message that I receive for

```bash
$ nvidia-smi -i 4
```

Unable to determine the device handle for GPU 0000:81:00.0: GPU is lost. Reboot the system to recover this GPU.

# BERT Multi-GPU TensorFlow and Horovod for prediction - no improvement

We have noticed a similar issue with PyTorch version of Transformer for NLP. This is the system config:
It is a Lambda blade server with 8 RTX 3090 GPUs, NVIDIA Driver Version: 460.56, CUDA Version: 11.2. Ubuntu 20.04 LTS, Dual AMD EPYC 7302 16-Core CPUs, 512GB of RAM.

# How to install libcudnn.8 Ubuntu 20.04 NVIDIA-SMI 460.56 Driver Version: 460.56 CUDA Version: 11.2

Got it: Followed this [python - "Could not load dynamic library 'libcudnn.so.8'" when running tensorflow on ubuntu 20.04 - Stack Overflow](https://stackoverflow.com/questions/66977227/could-not-load-dynamic-library-libcudnn-so-8-when-running-tensorflow-on-ubun)
Steps are:

```bash
wget https://developer.download.nvidia.com/compute/cuda/repos/ubuntu2004/x86_64/cuda-ubuntu2004.pin
sudo mv cuda-ubuntu2004.pin /etc/apt/preferences...
```

# Dist-upgrade does not upgrade to Cuda 11.1

Just a hunch, but I’m thinking there still might be some python*-*dev packages installed that are blocking this. Can you send me the output of

```bash
dpkg -l | grep python
```

```bash
dpkg -l | grep python | grep minimal
```

# After I upgraded (with several hoops and tries) to Ubuntu 20.04 Desktop

(from 18.04) via

```bash
sudo sed -i ‘/^AutomaticLoginEnable/ s/true/false/’ /etc/gdm3/custom.conf &&
sudo apt -y update && sudo apt -y dist-upgrade && sudo do-release-upgrade -d -f DistUpgradeViewNonInteractive
```

And then rebooted...

# How to replace the SSD in LambdaQuad?

I meant smartctl --all /dev/nvme0n1p2 says the SMART overall-health self-assessment test result: PASSED

I am in a situation where the main SSD in my LambdaQuad (4 GPU workstation server) is in danger of dying. I see the following error message once in a while: EXT4-fs error device nvme0n1p2 ext4-find-entry:1455 inode: nnnnnn System_Journal ID [nnn]: Failed to write entry… It is giving me warnings.

# GPU Utilization log/summary

nvidia-smi and related API can help us look at the current utilization of GPU Usage and Memory usage. Is there a tool or utility that can continuously log and give a summary of usage over a given time period like a day or a week? Thanks, -Karun.
