Skip to content

TutorialsAI & LLM

How to install vLLM on an NVIDIA GPU server with Docker

Install the NVIDIA driver and Container Toolkit, run the vLLM OpenAI-compatible server with Docker Compose, secure it with an API key and Caddy, tune memory.

  • Advanced
  • 45 min read
  • Updated

Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13

This guide is not available in your language yet, so it is shown in English.

On this page
  1. Prerequisites
  2. Step 1 — Install the NVIDIA driver
  3. Step 2 — Install the NVIDIA Container Toolkit
  4. Step 3 — Create the project directory and secrets
  5. Step 4 — Start vLLM with Docker Compose
  6. Step 5 — Test the OpenAI-compatible API
  7. Step 6 — Tune memory and use several GPUs
  8. Step 7 — Expose only the API through Caddy
  9. Alternative: install vLLM with uv and run it under systemd
  10. Back up and restore
  11. Update vLLM
  12. Troubleshooting
  13. CUDA out of memory
  14. could not select device driver with capabilities gpu
  15. CUDA driver is too old or a PTX toolchain error
  16. 401 or 403 when downloading a gated model
  17. The server hangs while downloading or starting
  18. Next steps

vLLM is a high-throughput inference and serving engine for large language models. It keeps many requests in flight on the GPU at once and exposes an OpenAI-compatible HTTP server, which makes it a common choice for production LLM APIs on NVIDIA hardware. This guide prepares a GPU server with NVIDIA's driver and Container Toolkit, runs the official vllm/vllm-openai image with Docker Compose on 127.0.0.1:8000, protects it with an API key and a Caddy proxy that exposes only the /v1 API, explains the memory and multi-GPU options, and shows an alternative installation with uv and systemd.

Prerequisites

ResourceMinimum (official)Suggested starting point
GPUNVIDIA, compute capability 7.5 or higherEnough VRAM for the model weights plus KV cache
Driver and CUDADefault builds use CUDA 12.9; CUDA 13 images need driver R580 or newerCurrent driver from NVIDIA's repository
Python (pip method)3.11 to 3.143.12, as in the official uv example
RAMNot publishedAt least as much system RAM as GPU memory
DiskNot published100 GB free for the image and model weights

The vLLM documentation does not publish RAM or disk minimums; the suggested values are a conservative starting point, not a benchmark. Check the size of a model's weight files on its Hugging Face page before you choose a GPU.

Step 1 — Install the NVIDIA driver

Install the driver from NVIDIA's network repository as described in NVIDIA's driver installation guide. The commands add the cuda-keyring package for your release and install the nvidia-open driver. On Debian, NVIDIA's guide also requires the contrib component. The guide enables it with add-apt-repository, which Debian 13 no longer ships, so first add contrib next to main in your APT sources yourself (the Components: line in /etc/apt/sources.list.d/debian.sources, or the deb lines in /etc/apt/sources.list, depending on which file your system uses):

Ubuntu

Bash
sudo apt update
sudo apt install linux-headers-$(uname -r)
distro=ubuntu$(. /etc/os-release && echo "$VERSION_ID" | tr -d .)
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install nvidia-open
sudo reboot

Debian

Bash
sudo apt update
sudo apt install linux-headers-$(uname -r)
distro=debian$(. /etc/os-release && echo "$VERSION_ID")
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt -V install nvidia-open
sudo reboot

After the reboot, run nvidia-smi. It lists every GPU, the driver version and, in the header, the highest CUDA version the driver supports. If your GPU needs an older driver branch, NVIDIA's guide explains how to pin one. If Secure Boot is enabled, the kernel module must be signed with a key the firmware trusts before it can load.

Step 2 — Install the NVIDIA Container Toolkit

The Container Toolkit lets Docker containers use the GPU. Add NVIDIA's repository, install the toolkit and configure Docker's runtime, following NVIDIA's installation guide:

Bash
sudo apt install -y --no-install-recommends ca-certificates curl gnupg2
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker

nvidia-ctk updates /etc/docker/daemon.json so that Docker can use the NVIDIA Container Runtime. Run NVIDIA's sample workload to confirm that containers see the GPU:

Bash
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi

You should see the same nvidia-smi table as on the host.

Step 3 — Create the project directory and secrets

Keep the Compose file, secrets and the Hugging Face cache together in /opt/vllm:

Bash
sudo mkdir -p /opt/vllm/hf-cache && sudo chown -R $USER:$USER /opt/vllm
cd /opt/vllm
echo "VLLM_API_KEY=$(openssl rand -hex 32)" > .env
echo "HF_TOKEN=" >> .env
chmod 600 .env

VLLM_API_KEY is read by the vLLM server as its API key. Leave HF_TOKEN empty for public models, or paste a Hugging Face access token after the = for gated models; for those, also request access on the model's Hugging Face page.

Step 4 — Start vLLM with Docker Compose

Create /opt/vllm/compose.yaml. The service follows the documented docker run command: all GPUs, the host IPC namespace (PyTorch shares memory between processes, which needs --ipc=host or a larger --shm-size), the Hugging Face cache mounted into the container and the API on port 8000, published on loopback only. This example serves the small Qwen/Qwen3-0.6B model used in the vLLM documentation:

YAML
services:
  vllm:
    image: vllm/vllm-openai:latest
    command: ["--model", "Qwen/Qwen3-0.6B", "--max-model-len", "8192"]
    environment:
      VLLM_API_KEY: "${VLLM_API_KEY}"
      HF_TOKEN: "${HF_TOKEN}"
    volumes:
      - ./hf-cache:/root/.cache/huggingface
    ports:
      - "127.0.0.1:8000:8000"
    ipc: host
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    restart: unless-stopped
Bash
docker compose up -d
docker compose logs -f vllm

The first start pulls a large image, downloads the model weights into /opt/vllm/hf-cache and prepares the GPU kernels, which can take several minutes. The server is ready when the log shows that the application started and Uvicorn is serving on port 8000.

Step 5 — Test the OpenAI-compatible API

Load the key into your shell and query the server. /health needs no key; /v1 endpoints do:

Bash
cd /opt/vllm
export VLLM_API_KEY=$(grep VLLM_API_KEY .env | cut -d= -f2)
curl -i http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $VLLM_API_KEY" \
  -d '{"model": "Qwen/Qwen3-0.6B", "messages": [{"role": "user", "content": "Say hello in five words"}]}'

/health returns 200 OK, /v1/models lists Qwen/Qwen3-0.6B, and the chat request returns a completion. A request to /v1/models without the header returns 401. Use --served-model-name if clients expect a different model name.

Step 6 — Tune memory and use several GPUs

vLLM reserves a share of each GPU's memory when it starts, loads the weights into it and uses the rest as KV cache for concurrent requests. Add these options to command in compose.yaml and run docker compose up -d to apply them:

OptionDefaultUse it to
--gpu-memory-utilization0.92Set the fraction of GPU memory this instance may use; lower it if other processes share the GPU
--max-model-lenFrom the model configCap the context length (prompt plus output); auto picks the largest length that fits
--max-num-seqsSet by vLLMLimit how many sequences are processed per iteration, which reduces memory use
--tensor-parallel-size, -tp1Split a model across several GPUs in one server
--quantization, -qFrom the model configSelect a quantization method for smaller weights
--enforce-eageroffDisable CUDA graphs, which saves some GPU memory

For example, to spread a larger model across two GPUs, use command: ["--model", "org/model-name", "--tensor-parallel-size", "2"]. To dedicate specific GPUs to one instance, set CUDA_VISIBLE_DEVICES in environment, for example CUDA_VISIBLE_DEVICES: "0,1". Multi-node deployments need an isolated network between the nodes, because vLLM's internal communication is not secured.

Step 7 — Expose only the API through Caddy

The vLLM security documentation recommends a reverse proxy that allows only the endpoints you want to expose, because --api-key protects /v1, /v2, /inference and /cohere but not endpoints such as /tokenize, /pooling or /health. This Caddy site block forwards /v1/* and answers everything else with 404:

Caddyfile
llm.example.com {
    handle /v1/* {
        reverse_proxy 127.0.0.1:8000
    }
    handle {
        respond 404
    }
}
Bash
sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable

Clients now use https://llm.example.com/v1 as the base URL and the value of VLLM_API_KEY as their API key. Streaming responses pass through Caddy without extra settings.

Alternative: install vLLM with uv and run it under systemd

If you cannot use Docker, the documentation recommends installing vLLM with uv in a fresh virtual environment. Install uv with its official installer after reviewing it, then create the environment in /opt/vllm with a managed Python 3.12 stored inside the project directory, so the service user can read it:

Bash
sudo apt install build-essential
curl -LsSf https://astral.sh/uv/install.sh -o uv-install.sh
less uv-install.sh
sh uv-install.sh
source $HOME/.local/bin/env
sudo mkdir -p /opt/vllm && sudo chown $USER:$USER /opt/vllm
cd /opt/vllm
export UV_PYTHON_INSTALL_DIR=/opt/vllm/python
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
python -c "import vllm; print(vllm.__version__)"

--torch-backend=auto selects the PyTorch build that matches your installed driver. Next, create a service user and a root-only environment file:

Bash
sudo useradd --system --create-home --home-dir /var/lib/vllm --shell /usr/sbin/nologin vllm
echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/vllm.env > /dev/null
echo "HF_HOME=/var/lib/vllm/huggingface" | sudo tee -a /etc/vllm.env > /dev/null
sudo chmod 600 /etc/vllm.env

Create /etc/systemd/system/vllm.service, then run sudo systemctl daemon-reload and sudo systemctl enable --now vllm:

INI
[Unit]
Description=vLLM OpenAI-compatible server
After=network-online.target
Wants=network-online.target

[Service]
User=vllm
Group=vllm
EnvironmentFile=/etc/vllm.env
ExecStart=/opt/vllm/.venv/bin/vllm serve Qwen/Qwen3-0.6B --host 127.0.0.1 --port 8000 --max-model-len 8192
Restart=on-failure
RestartSec=10

[Install]
WantedBy=multi-user.target

The --host 127.0.0.1 option keeps the server on loopback. Follow the start-up with journalctl -u vllm -f, test it as in Step 5 and use the same Caddy block. Add HF_TOKEN= to /etc/vllm.env for gated models.

Back up and restore

vLLM keeps no user data. Back up the configuration: compose.yaml and .env (or /etc/vllm.env and the unit file). The Hugging Face cache can be downloaded again; include it only if downloads are slow or the model is private.

Bash
sudo mkdir -p /opt/backups
sudo tar czf /opt/backups/vllm-config-$(date +%F).tar.gz /opt/vllm/compose.yaml /opt/vllm/.env

To restore, prepare a server with Steps 1 and 2, extract the archive with sudo tar xzf /opt/backups/vllm-config-YYYY-MM-DD.tar.gz -C /, fix the ownership of /opt/vllm and run docker compose up -d. Copy the archive off the server; it contains your API key.

Update vLLM

Read the release notes on the vLLM releases page first: options and defaults change between versions. For Docker, change the pinned tag (or keep latest), pull and recreate:

Bash
cd /opt/vllm
docker compose pull
docker compose up -d
docker image prune

For the uv installation, the documentation recommends a fresh environment rather than upgrading in place, because compiled kernels are tied to specific CUDA and PyTorch versions. Create a new virtual environment next to the old one, install vLLM into it, point ExecStart at the new path and restart the service; keep the old environment until the new one works. Driver updates arrive through apt upgrade from NVIDIA's repository and need a reboot.

Troubleshooting

CUDA out of memory

The weights plus KV cache do not fit. Lower --max-model-len or --max-num-seqs, use a quantized version of the model, split it with --tensor-parallel-size, or check with nvidia-smi that no other process holds GPU memory. If other workloads share the GPU, lower --gpu-memory-utilization. The vLLM memory guide also suggests --enforce-eager to skip CUDA graph memory.

could not select device driver with capabilities gpu

Docker does not know the NVIDIA runtime. Repeat Step 2, especially sudo nvidia-ctk runtime configure --runtime=docker and the Docker restart, and test with the sample workload.

CUDA driver is too old or a PTX toolchain error

The driver is older than the CUDA version of the image or wheel. Update the driver from NVIDIA's repository and reboot. For some datacenter GPUs the vLLM documentation offers a compatibility mode: add VLLM_ENABLE_CUDA_COMPATIBILITY: "1" to the container environment.

401 or 403 when downloading a gated model

HF_TOKEN is missing or invalid, or your Hugging Face account has not been granted access to the model yet. Fix the token in .env, request access on the model's Hugging Face page, wait until it is granted and run docker compose up -d.

The server hangs while downloading or starting

Download the model separately with the Hugging Face hf command line tool into the cache directory and start vLLM again; this shows whether the download is the problem. For more output, set VLLM_LOGGING_LEVEL: "DEBUG" in the environment, and remove it again when you are done.

Next steps

Frequently asked questions

Which GPUs does vLLM support?

vLLM's CUDA builds need an NVIDIA GPU with compute capability 7.5 or higher, for example T4, RTX 20 series and newer, A100, L4, H100 or B200. It runs on Linux only; there is no native Windows support.

Is the vLLM API key enough to secure the server?

No. The vLLM documentation states that --api-key only authenticates endpoints under /v1, /v2, /inference and /cohere, while other endpoints on the same server stay open. Keep vLLM on 127.0.0.1 and let a reverse proxy expose only the paths you need, as this guide does with Caddy.

How much GPU memory does a model need?

The model weights must fit in GPU memory, and vLLM uses the rest of its share for the KV cache. By default an instance may use 92 percent of each GPU’s memory (--gpu-memory-utilization 0.92). Larger models need quantized weights or tensor parallelism across several GPUs.

How do I serve a gated model such as Llama?

Request access on the model’s Hugging Face page with your account, create an access token and put it in HF_TOKEN in the .env file. Once access is granted, vLLM downloads the weights into the Hugging Face cache volume.

Should I use Docker or pip to install vLLM?

The official vllm/vllm-openai image bundles a matching CUDA and PyTorch stack and is the simplest way to run the server. The pip or uv install in a virtual environment suits custom setups; the docs recommend a fresh environment because compiled kernels are tied to specific CUDA and PyTorch versions.

Sources

מחולל סיסמאות

Please confirm