Deploying Llama 3.2 3B in a Kubernetes (K8s) cluster - Lambda Docs

Deploying Llama 3.2 3B in a Kubernetes (K8s) cluster

Introduction

In this tutorial, you'll:

  1. Stand up a single-node Kubernetes cluster on an on-demand instance using K3s.
  2. Install the NVIDIA GPU Operator so your cluster can use your instance's GPUs.
  3. Deploy Ollama in your cluster to serve the Llama 3.2 3B model.
  4. Install the Ollama client.
  5. Interact with the Llama 3.2 3B model.

Note
You don't need a Kubernetes cluster to run Ollama and serve the Llama 3.2 3B model. Part of this tutorial is to demonstrate that it's possible to stand up a Kubernetes cluster on on-demand instances.

Stand up a single-node Kubernetes cluster

  1. If you haven't already, use the console or Cloud API to launch an instance. Then, SSH into your instance.

  2. Install K3s (Kubernetes) by running:

   curl -sfL https://get.k3s.io | K3S_KUBECONFIG_MODE=644 sh -s - --default-runtime=nvidia
  1. Verify that your Kubernetes cluster is ready by running:
   k3s kubectl get nodes

You should see output similar to:

NAME              STATUS   ROLES                  AGE    VERSION
104-171-203-164   Ready    control-plane,master   100s   v1.30.5+k3s1
  1. Install socat by running:
   sudo apt -y install socat

socat is needed to enable port forwarding in a later step.

Install the NVIDIA GPU Operator

  1. Install the NVIDIA GPU Operator in your Kubernetes cluster by running:
   cat <<EOF | k3s kubectl apply -f -
   apiVersion: v1
   kind: Namespace
   metadata:
        name: gpu-operator
      ---
   apiVersion: helm.cattle.io/v1
   kind: HelmChart
   metadata:
        name: gpu-operator
        namespace: gpu-operator
   spec:
        repo: https://helm.ngc.nvidia.com/nvidia
        chart: gpu-operator
        targetNamespace: gpu-operator
   EOF
  1. In a few minutes, verify that your instance's GPUs are detected by your cluster by running:
   k3s kubectl describe nodes | grep nvidia.com

You should see output similar to:

nvidia.com/cuda.driver-version.full=535.129.03
nvidia.com/cuda.driver-version.major=535
nvidia.com/cuda.driver-version.minor=129
nvidia.com/cuda.driver.version.revision=03
nvidia.com/cuda.runtime-version.full=12.2
nvidia.com/cuda.runtime-version.major=12
nvidia.com/cuda.runtime-version.minor=2
nvidia.com/gpu.count=8
nvidia.com/gpu.product=Tesla-V100-SXM2-16GB

nvidia.com/gpu.count=8 indicates that your cluster detects 8 GPUs.

nvidia.com/gpu.product=Tesla-V100-SXM2-16GB indicates that the detected GPUs are Tesla V100 SXM2 16GB GPUs.

Note
In this tutorial, Ollama will only use 1 GPU.

Deploy Ollama in your Kubernetes cluster

  1. Start an Ollama server in your Kubernetes cluster by running:
   cat <<EOF | k3s kubectl apply -f -
   apiVersion: v1
   kind: Namespace
   metadata:
        name: ollama
      ---
   apiVersion: apps/v1
   kind: Deployment
   metadata:
        name: ollama
        namespace: ollama
   spec:
        strategy:
          type: Recreate
        selector:
          matchLabels:
            name: ollama
        template:
          metadata:
            labels:
              name: ollama
          spec:
            containers:
         - name: ollama
           image: ollama/ollama:latest
           env:
           - name: PATH
             value: /usr/local/nvidia/bin:/usr/local/cuda/bin:/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin
           - name: LD_LIBRARY_PATH
             value: /usr/local/nvidia/lib:/usr/local/nvidia/lib64
           - name: NVIDIA_DRIVER_CAPABILITIES
             value: compute,utility
           ports:
           - name: http
             containerPort: 11434
             protocol: TCP
           resources:
             limits:
               nvidia.com/gpu: 1
         tolerations:
         - key: nvidia.com/gpu
           operator: Exists
           effect: NoSchedule
   ---
   apiVersion: v1
   kind: Service
   metadata:
   name: ollama
   namespace: ollama
   spec:
   type: ClusterIP
   selector:
       name: ollama
   ports:
     - port: 80
       name: http
       targetPort: http
       protocol: TCP
   EOF
  1. In a few minutes, run the following command to verify that the Ollama server is accepting connections and is using a GPU:
   kubectl logs -n ollama -l name=ollama

You should see output similar to:

2024/09/27 18:51:55 routes.go:1153: INFO server config ...

The last line in the example output above shows that Ollama is using a single Tesla V100-SXM2-16GB GPU.

  1. Start a tmux session by running:
   tmux

Then, run the following command to make Ollama accessible from outside of your Kubernetes cluster:

   k3s kubectl -n ollama port-forward service/ollama 11434:80

You should see output similar to:

Forwarding from 127.0.0.1:11434 -> 11434

Install the Ollama client

Download and install the Ollama client by running:

   curl -L https://ollama.com/download/ollama-linux-amd64.tgz -o ollama-linux-amd64.tgz
   sudo tar -C /usr -xzf ollama-linux-amd64.tgz

Serve and interact with the Llama 3.2 3B model

  1. Serve the Llama 3.2 3B model using Ollama by running:
   ollama run llama3.2:3b-instruct-fp16

You can interact with the model once you see the following prompt:

>>> Send a message (/? for help)
  1. Test the model by entering a prompt, for example:
    What is machine learning? Explain like I'm five.
    

You should see output similar to:

MACHINE LEARNING IS SO COOL!
...

That's basically what machine learning is: teaching a computer to recognize patterns and make decisions on its own, just like the robot did with the toys!