Introduction - Lambda Docs

Using Lambda's Managed Slurm

Ready to jump in?

The quick start takes you from SSH to your first multi-node GPU job in minutes.

Introduction to Slurm

Slurm is a widely used open-source workload manager optimized for high-performance computing (HPC) and machine learning (ML) workloads. Slurm allows administrators to create user accounts with controlled access, enabling individual users to submit, monitor, and manage their workloads.

Slurm automatically schedules workloads, maximizing cluster utilization while preventing resource contention.

Lambda's Slurm

The table below summarizes the key differences between Lambda's Managed and Unmanaged Slurm deployments:

Feature Managed Slurm ✓ Unmanaged Slurm ✗
Slurm-managed compute node access ✓ ✗ (out-of-band access allowed)
User sudo/root privileges ✗ ✓
Lambda monitors Slurm daemons ✓ ✗ (customer is responsible)
Lambda applies patches and upgrades ✓ (on request) ✗ (customer is responsible)
Slurm support with SLAs ✓ ✗
Lambda Slurm configuration ✓ ✓
Slurm configured for high availability ✓ ✓
Shared /home across all nodes ✓ ✓
Shared /data across all nodes ✓ ✓

Managed Slurm

When Lambda's Managed Slurm (MSlurm) is deployed:

Unmanaged Slurm

In contrast, with Unmanaged Slurm:

Warning

Workloads that run outside of Slurm might interfere with the resources managed by Slurm. Additionally, users with administrator access can make changes that render the cluster unrecoverable. In such cases, Lambda might need to "repave" the cluster, fully wiping and reinstalling the system.

Shared features

Both Managed and Unmanaged Slurm configurations include:

Note

It's recommended to stage data, such as datasets and models, on local storage before running a job. Accessing files directly from shared storage during a job can lead to degraded performance due to I/O bottlenecks.

Accessing the MSlurm cluster

Your first stop is the Slurm console, opened from the Lambda Cloud console: see the whole cluster, add accounts, and watch jobs without touching a terminal. For day-to-day work, connect over SSH as described below.

The MSlurm cluster is initially configured with a single user account: ubuntu. This account is pre-configured with the SSH key provided and is used to administer the cluster. Additional user accounts can be created from the Slurm console or from the ubuntu account. See Creating and removing MSlurm users.

To access the MSlurm cluster as the ubuntu user, SSH into the login node. Replace <LOGIN-NODE-IP> with the IP address of -login-001, available in the Lambda Cloud console:

ssh ubuntu@<LOGIN-NODE-IP>

Other users access the MSlurm cluster in the same way, by using SSH to log into the login node. Replace <USERNAME> with the appropriate username:

ssh <USERNAME>@<LOGIN-NODE-IP>

Creating and removing MSlurm users

In the MSlurm cluster, user accounts and groups control both system access and job submission permissions. LDAP provides consistent user and group management across all nodes by acting as a centralized directory.

Like standard Slurm installations, MSlurm doesn't maintain its own user database. Instead, it relies on the underlying system's authentication and group management.

Lambda provides two easy ways to manage accounts: the Slurm console in your browser, and the suser tool on the login node.

Managing users from the Slurm console

The Slurm console is the fastest way to manage users: its Cluster Users app lists every account alongside its status, Slurm accounts, and SSH keys. Adding a user is one short form:

Everything except the username is optional and can be filled in later, so adding a teammate takes seconds.

Creating a new user from the command line

To create a new user using suser:

  1. SSH into the MSlurm login node using the ubuntu account. Replace <LOGIN-NODE-IP> with the IP address of the login node (-login-001):
ssh ubuntu@<LOGIN-NODE-IP>
  1. Create the user. Replace <USERNAME> with the desired username, and <SSH-KEY> with either the path to the user's SSH public key file or the public key string itself:
suser add <USERNAME> --key <SSH-KEY>

After the command completes, the message User <USERNAME> successfully added will confirm the user was created.

Removing a user from the command line

To remove a user using suser:

  1. SSH into the MSlurm login node using the ubuntu account. Replace <LOGIN-NODE-IP> with the IP address of the login node (-login-001):
ssh ubuntu@<LOGIN-NODE-IP>
  1. Remove the user. Replace <USERNAME> with the actual username:
suser remove <USERNAME>

After the command completes, the message User <USERNAME> successfully removed will confirm the user was removed.

suser never deletes home directories.

suser --help shows you all available options.

Running jobs on the MSlurm cluster

Jobs are submitted to the MSlurm cluster using the sbatch, srun, and salloc commands:

The MSlurm cluster supports Pyxis and Enroot, enabling srun to run containers, including those based on Docker images, on compute nodes.

Below are examples of using sbatch, srun, and salloc. Two sbatch examples are included: one that runs nvidia-smi -L on compute nodes and displays its output, and another that evaluates a language model's ability to solve multiplication problems. All examples print the hostnames of the compute nodes where nvidia-smi -L was executed.

The examples should be run on the MSlurm cluster login node.

Using sbatch to run nvidia-smi -L

  1. Create a file named nvidia_smi_batch.sh containing the following:
#!/bin/bash
#SBATCH --nodes=2
#SBATCH --gpus=2
#SBATCH --ntasks=2
#SBATCH --ntasks-per-node=1
#SBATCH --output="sbatch_output_direct_%x_%j.out"
#SBATCH --error="sbatch_output_direct_%x_%j.err"
#SBATCH --time=00:01:00

echo "Job ID: $SLURM_JOB_ID"
echo "Running on nodes: $SLURM_NODELIST"
echo
srun --ntasks=$SLURM_NTASKS nvidia-smi -L
  1. Submit the job using sbatch:
sbatch nvidia_smi_batch.sh

This command submits the job and performs the following steps:

  1. Requests cluster resources:
    • --nodes=2: Reserves 2 compute nodes.
    • --gpus=2: Requests a total of 2 GPUs across all nodes.
    • --ntasks=2: Runs 2 parallel tasks in total.
    • --ntasks-per-node=1: Assigns 1 task per node (2 tasks across 2 nodes).
  2. Configures job output:
    • --output="sbatch_output_direct_%x_%j.out": Saves standard output to a file named with the job name (%x) and job ID (%j).
    • --error="sbatch_output_direct_%x_%j.err": Saves standard error to a similar file.
  3. Sets a job time limit:
    • --time=00:01:00: Limits the job runtime to 1 minute.
  4. Prints the job information:
    • echo "Job ID: $SLURM_JOB_ID": Displays the assigned job ID.
    • echo "Running on nodes: $SLURM_NODELIST": Displays the list of allocated nodes.
  5. Runs nvidia-smi -L:
    • srun --ntasks=$SLURM_NTASKS nvidia-smi -L: Runs nvidia-smi -L on all tasks to list visible GPUs.

After the job completes, two files are created:

Contains the job ID, allocated nodes, and the output of nvidia-smi -L from each task.

Contains any error messages. This file is usually empty unless something went wrong.

Using sbatch to evaluate a large language model (LLM)

As an additional example, a Slurm batch job can be used to evaluate how well an LLM solves basic multiplication problems:

  1. Download the Python script and the Slurm batch script:
curl -sSLO https://docs.lambda.ai/assets/code/eval_multiplication.py
curl -sSLO https://docs.lambda.ai/assets/code/run_eval.sh

Both scripts are annotated with comments explaining their structure and purpose.

  1. Submit the job using a Hugging Face model ID:
sbatch run_eval.sh deepseek-ai/DeepSeek-R1-Distill-Llama-70B

The model ID can be replaced with any other compatible model.

  1. To follow the job's progress in real time:
tail -F eval_results.out

This log shows when the model is loading, prompts are being processed, and sampling is running.

  1. After the job completes, the accuracy is saved to a file in the accuracies/ directory. To view it:
cat accuracies/deepseek-ai_DeepSeek-R1-Distill-Llama-70B.txt

The filename matches the model ID with slashes replaced by underscores.

Using srun to run nvidia-smi -L

Below are two methods for running nvidia-smi -L with srun:

  1. Direct execution on compute nodes: Runs nvidia-smi -L directly on the assigned nodes.
  2. Execution inside containers: Runs nvidia-smi -L within a containerized environment on the compute nodes.

Direct execution on compute nodes

srun --gpus=2 --nodes=2 --ntasks-per-node=1 \
     --output="srun_output_direct_%N.txt" \
     bash -c 'printf "\n===== Node: $(hostname) =====\n"; nvidia-smi -L'

This command runs nvidia-smi -L directly on two compute nodes and saves the output in separate text files. The filenames are based on the hostnames of the respective nodes, for example:

Each file contains the nvidia-smi -L output from its corresponding compute node.

Execution inside containers

srun --gpus=2 --nodes=2 --ntasks-per-node=1 \
     --output="srun_output_container_%N.txt" \
     --container-image=nvidia/cuda:12.8.1-runtime-ubuntu22.04 \
     bash -c 'printf "\n===== Node: $(hostname) =====\n"; nvidia-smi -L'

This command performs the same task as above but runs nvidia-smi -L inside an NVIDIA CUDA container instead of directly on the compute nodes. The output is saved in separate files, such as:

Each file contains the nvidia-smi -L output from its respective node while running within a containerized environment.

Using salloc to run nvidia-smi -L

Unlike srun and sbatch, salloc doesn't run a single specified task. Instead, it allocates resources and opens an interactive shell directly on the allocated compute node. Run commands as you like, then exit the shell to release the allocation.

In this example, one node with two GPUs is requested, rather than two nodes with one GPU each as shown in the earlier sbatch and srun examples. This difference is reflected in the nvidia-smi -L output. The output appears directly in the terminal instead of being saved to a file.

  1. Allocate one node with two GPUs and start an interactive shell on the allocated node:
salloc --gpus=2 --nodes=1
  1. Print the node's hostname and run nvidia-smi -L:
printf "\n===== Node: $(hostname) =====\n"; nvidia-smi -L
  1. Exit the interactive shell and release the allocated resources by pressing Ctrl + D.

Managing software using Lmod

Both Managed and Unmanaged Slurm include the Lmod module system by default. You can use Lmod's module command-line tool to dynamically load and unload software modules. When you load a module, Lmod updates your shell environment so the selected software is available.

When you load a module on the login node, that environment is exported to your Slurm jobs automatically. For example, if you run module load uv on the login node and then submit a job using srun or sbatch, uv will be available on the compute nodes without needing to reload it.

Common Lmod commands:

Command Description
module avail List all modules available to be loaded.
module spider List and describe every module available on the cluster, including those hidden by the current hierarchy.
module load <MODULE> Load a specific module into your current environment.
module list Show all modules currently loaded in your session.
module unload <MODULE> Remove a specific module from your environment.
module purge Unload all modules to start with a clean environment.

Tip

When submitting jobs using sbatch, include the module load commands inside your batch script. This ensures the compute nodes have the correct environment configured before your code executes.

Next steps