Skip to main content

Setting up NVIDIA and CUDA Environment

This document describes how to use ocboot's setup-ai-env command to configure NVIDIA drivers, CUDA, containerd, and other AI runtime environments on an existing cluster or standalone machine, without deploying or modifying Cloudpods cloud management services. This applies to scenarios that require GPU-dependent AI container applications such as vLLM and SGLang.

Notes
  • If the target environment does not have an Nvidia GPU, or you do not need to run GPU-dependent container AI applications like vLLM or SGLang, you can skip this document.
  • Without the Nvidia runtime environment configured, you can still run AI container applications that do not depend on GPU, such as OpenClaw and Dify.

Download Driver Files​

Download and prepare the NVIDIA driver and CUDA installation packages (example: CUDA 13.0.3; adjust for your GPU and compatibility). Use .run packages:

Transfer the packages to the ocboot code directory on the machine that runs ocboot.sh.

warning

The NVIDIA driver and CUDA installer packages must be placed in the ocboot code directory. ocboot.sh runs deployment inside a buildah container; other host paths are not mapped in. Inside the container the path is /ocboot/<filename>.

If the packages are on another machine, copy them to the ocboot code directory on the target host (replace with your actual path):

# Using rsync (recommended)
rsync -avP /path/to/nvidia/NVIDIA-Linux-x86_64-580.173.02.run target_host:/path/to/ocboot/
rsync -avP /path/to/cuda/cuda_13.0.3_580.126.20_linux.run target_host:/path/to/ocboot/

Command Format​

Run ocboot.sh setup-ai-env to set up the nvidia container runtime on the target machine.

./ocboot.sh setup-ai-env <target_host1> [target_host2 ...] \
--nvidia-driver-installer-path ./<driver_file>.run \
--cuda-installer-path ./<cuda_file>.run \
[--gpu-device-virtual-number 2] [--user USER] [--key-file KEY] [--port PORT]

Examples​

Run from the ocboot code directory (<target_host> is the target IP; for a single-node deploy, use the local IP):

The target host will reboot automatically
  • After setup finishes, the target host usually reboots once to load the new driver and GRUB configuration. This is expected.
  • If nouveau is still using the GPU, the host may reboot once earlier before installing the NVIDIA driver.
  • Do not interrupt the process during reboot. If installation is incomplete after reboot, re-run the same setup-ai-env command until all steps finish.
# Basic usage
./ocboot.sh setup-ai-env <target_host> \
--nvidia-driver-installer-path ./NVIDIA-Linux-x86_64-580.173.02.run \
--cuda-installer-path ./cuda_13.0.3_580.126.20_linux.run

# Specify GPU share virtual number (NVIDIA_GPU_SHARE) and SSH options; omit to use HAMi by default
./ocboot.sh setup-ai-env <target_host> \
--nvidia-driver-installer-path ./NVIDIA-Linux-x86_64-580.173.02.run \
--cuda-installer-path ./cuda_13.0.3_580.126.20_linux.run \
--gpu-device-virtual-number 2 \
--user admin \
--port 2222

Parameter Reference​

ParameterRequiredDefaultDescription
--nvidia-driver-installer-pathYes-Full path to the NVIDIA driver installation package.
--cuda-installer-pathYes-Full path to the CUDA installation package.
--gpu-device-virtual-numberNonone (HAMi by default)When set, creates NVIDIA_GPU_SHARE virtual devices; when omitted, GPUs use HAMi.
--user / -uNorootSSH username.
--key-file / -kNo-SSH private key file path.
--port / -pNo22SSH port.

Installation Steps and Workflow​

ocboot will perform the following steps (among others) on the target host:

  1. Check OS support and local installation files (if paths are specified)
  2. Configure GRUB (add nvidia-drm.modeset=1)
  3. Install kernel headers and development packages
  4. Clean up vfio-related configurations (if present)
  5. Install NVIDIA driver (if --nvidia-driver-installer-path is provided)
  6. Install CUDA environment (if --cuda-installer-path is provided)
  7. Install NVIDIA Container Toolkit
  8. Configure containerd runtime
  9. Configure host device mappings (only when --gpu-device-virtual-number is set: discover /dev/dri/renderD* and generate NVIDIA_GPU_SHARE config; otherwise HAMi is used by default)
  10. Reboot if needed, then verify the installation

Verify the Driver​

On the target host:

nvidia-smi

You should see the GPU model, driver version, and memory information. Confirm that the driver version matches the CUDA version.

Important Notes​

  • Ensure the target host has sufficient disk space and network connectivity to download dependencies such as the NVIDIA Container Toolkit.
  • Prepare packages on the machine running ocboot, and transfer them to the ocboot code directory on the target before running the playbook.
  • Ensure passwordless SSH login between the machine running ocboot and the target host (or use parameters such as --key-file).

FAQ​

No GPU listed or GPU not visible in the console after deployment?​

  1. On the node, run nvidia-smi to confirm the driver works.
  2. Confirm the driver and CUDA versions match.
  3. Confirm the host is enabled (Compute → Infrastructure → Hosts).
  4. Check Compute → Infrastructure → Passthrough Devices for reported GPUs.
  5. If needed, re-run the same setup-ai-env command.

If you used run.py ai without --nvidia-driver-installer-path / --cuda-installer-path, install the driver and CUDA on the target first, or configure them with setup-ai-env as described in this document.

Is it normal for the target host to automatically reboot during installation?​

Yes. ocboot usually reboots once after all install steps finish, to load the new driver or GRUB configuration. If nouveau is still using the GPU, it may also reboot once before installing the driver. Do not interrupt the process. If deployment is incomplete after reboot, re-run the same setup-ai-env command to continue.

How to check if the platform has GPUs available?​

If configuration succeeds, go to Compute → Infrastructure → Passthrough Devices to see detected GPUs. When --gpu-device-virtual-number is omitted, the sharing mode is HAMI.