How to install vLLM on an NVIDIA GPU server with Docker
Install the NVIDIA driver and Container Toolkit, run the vLLM OpenAI-compatible server with Docker Compose, secure it with an API key and Caddy, tune memory.
- Advanced
- 45 min read
- Updated
Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13
This guide is not available in your language yet, so it is shown in English.
On this page
- Prerequisites
- Step 1 — Install the NVIDIA driver
- Step 2 — Install the NVIDIA Container Toolkit
- Step 3 — Create the project directory and secrets
- Step 4 — Start vLLM with Docker Compose
- Step 5 — Test the OpenAI-compatible API
- Step 6 — Tune memory and use several GPUs
- Step 7 — Expose only the API through Caddy
- Alternative: install vLLM with uv and run it under systemd
- Back up and restore
- Update vLLM
- Troubleshooting
- CUDA out of memory
- could not select device driver with capabilities gpu
- CUDA driver is too old or a PTX toolchain error
- 401 or 403 when downloading a gated model
- The server hangs while downloading or starting
- Next steps
vLLM is a high-throughput inference and serving engine for large language models. It keeps many requests in flight on the GPU at once and exposes an OpenAI-compatible HTTP server, which makes it a common choice for production LLM APIs on NVIDIA hardware. This guide prepares a GPU server with NVIDIA's driver and Container Toolkit, runs the official vllm/vllm-openai image with Docker Compose on 127.0.0.1:8000, protects it with an API key and a Caddy proxy that exposes only the /v1 API, explains the memory and multi-GPU options, and shows an alternative installation with uv and systemd.
Prerequisites
- A dedicated server with a supported NVIDIA GPU; see GPU servers. vLLM's CUDA builds need compute capability 7.5 or newer.
- Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12 or Debian 13 on x86_64. NVIDIA's driver guide covers all four releases.
- A non-root user with
sudorights; see Secure a new Linux server and Set up SSH keys. - Docker Engine with the Compose plugin from Install Docker on Ubuntu or Install Docker on Debian.
- A Hugging Face account and access token if you want to serve gated models.
- A domain such as
llm.example.compointing at the server, and Caddy from Caddy as a reverse proxy.
| Resource | Minimum (official) | Suggested starting point |
|---|---|---|
| GPU | NVIDIA, compute capability 7.5 or higher | Enough VRAM for the model weights plus KV cache |
| Driver and CUDA | Default builds use CUDA 12.9; CUDA 13 images need driver R580 or newer | Current driver from NVIDIA's repository |
| Python (pip method) | 3.11 to 3.14 | 3.12, as in the official uv example |
| RAM | Not published | At least as much system RAM as GPU memory |
| Disk | Not published | 100 GB free for the image and model weights |
The vLLM documentation does not publish RAM or disk minimums; the suggested values are a conservative starting point, not a benchmark. Check the size of a model's weight files on its Hugging Face page before you choose a GPU.
Step 1 — Install the NVIDIA driver
Install the driver from NVIDIA's network repository as described in NVIDIA's driver installation guide. The commands add the cuda-keyring package for your release and install the nvidia-open driver. On Debian, NVIDIA's guide also requires the contrib component. The guide enables it with add-apt-repository, which Debian 13 no longer ships, so first add contrib next to main in your APT sources yourself (the Components: line in /etc/apt/sources.list.d/debian.sources, or the deb lines in /etc/apt/sources.list, depending on which file your system uses):
Ubuntu
sudo apt update
sudo apt install linux-headers-$(uname -r)
distro=ubuntu$(. /etc/os-release && echo "$VERSION_ID" | tr -d .)
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install nvidia-open
sudo rebootDebian
sudo apt update
sudo apt install linux-headers-$(uname -r)
distro=debian$(. /etc/os-release && echo "$VERSION_ID")
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt -V install nvidia-open
sudo rebootAfter the reboot, run nvidia-smi. It lists every GPU, the driver version and, in the header, the highest CUDA version the driver supports. If your GPU needs an older driver branch, NVIDIA's guide explains how to pin one. If Secure Boot is enabled, the kernel module must be signed with a key the firmware trusts before it can load.
Step 2 — Install the NVIDIA Container Toolkit
The Container Toolkit lets Docker containers use the GPU. Add NVIDIA's repository, install the toolkit and configure Docker's runtime, following NVIDIA's installation guide:
sudo apt install -y --no-install-recommends ca-certificates curl gnupg2
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart dockernvidia-ctk updates /etc/docker/daemon.json so that Docker can use the NVIDIA Container Runtime. Run NVIDIA's sample workload to confirm that containers see the GPU:
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smiYou should see the same nvidia-smi table as on the host.
Step 3 — Create the project directory and secrets
Keep the Compose file, secrets and the Hugging Face cache together in /opt/vllm:
sudo mkdir -p /opt/vllm/hf-cache && sudo chown -R $USER:$USER /opt/vllm
cd /opt/vllm
echo "VLLM_API_KEY=$(openssl rand -hex 32)" > .env
echo "HF_TOKEN=" >> .env
chmod 600 .envVLLM_API_KEY is read by the vLLM server as its API key. Leave HF_TOKEN empty for public models, or paste a Hugging Face access token after the = for gated models; for those, also request access on the model's Hugging Face page.
Step 4 — Start vLLM with Docker Compose
Create /opt/vllm/compose.yaml. The service follows the documented docker run command: all GPUs, the host IPC namespace (PyTorch shares memory between processes, which needs --ipc=host or a larger --shm-size), the Hugging Face cache mounted into the container and the API on port 8000, published on loopback only. This example serves the small Qwen/Qwen3-0.6B model used in the vLLM documentation:
services:
vllm:
image: vllm/vllm-openai:latest
command: ["--model", "Qwen/Qwen3-0.6B", "--max-model-len", "8192"]
environment:
VLLM_API_KEY: "${VLLM_API_KEY}"
HF_TOKEN: "${HF_TOKEN}"
volumes:
- ./hf-cache:/root/.cache/huggingface
ports:
- "127.0.0.1:8000:8000"
ipc: host
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]
restart: unless-stoppeddocker compose up -d
docker compose logs -f vllmThe first start pulls a large image, downloads the model weights into /opt/vllm/hf-cache and prepares the GPU kernels, which can take several minutes. The server is ready when the log shows that the application started and Uvicorn is serving on port 8000.
Step 5 — Test the OpenAI-compatible API
Load the key into your shell and query the server. /health needs no key; /v1 endpoints do:
cd /opt/vllm
export VLLM_API_KEY=$(grep VLLM_API_KEY .env | cut -d= -f2)
curl -i http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl http://127.0.0.1:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer $VLLM_API_KEY" \
-d '{"model": "Qwen/Qwen3-0.6B", "messages": [{"role": "user", "content": "Say hello in five words"}]}'/health returns 200 OK, /v1/models lists Qwen/Qwen3-0.6B, and the chat request returns a completion. A request to /v1/models without the header returns 401. Use --served-model-name if clients expect a different model name.
Step 6 — Tune memory and use several GPUs
vLLM reserves a share of each GPU's memory when it starts, loads the weights into it and uses the rest as KV cache for concurrent requests. Add these options to command in compose.yaml and run docker compose up -d to apply them:
| Option | Default | Use it to |
|---|---|---|
--gpu-memory-utilization | 0.92 | Set the fraction of GPU memory this instance may use; lower it if other processes share the GPU |
--max-model-len | From the model config | Cap the context length (prompt plus output); auto picks the largest length that fits |
--max-num-seqs | Set by vLLM | Limit how many sequences are processed per iteration, which reduces memory use |
--tensor-parallel-size, -tp | 1 | Split a model across several GPUs in one server |
--quantization, -q | From the model config | Select a quantization method for smaller weights |
--enforce-eager | off | Disable CUDA graphs, which saves some GPU memory |
For example, to spread a larger model across two GPUs, use command: ["--model", "org/model-name", "--tensor-parallel-size", "2"]. To dedicate specific GPUs to one instance, set CUDA_VISIBLE_DEVICES in environment, for example CUDA_VISIBLE_DEVICES: "0,1". Multi-node deployments need an isolated network between the nodes, because vLLM's internal communication is not secured.
Step 7 — Expose only the API through Caddy
The vLLM security documentation recommends a reverse proxy that allows only the endpoints you want to expose, because --api-key protects /v1, /v2, /inference and /cohere but not endpoints such as /tokenize, /pooling or /health. This Caddy site block forwards /v1/* and answers everything else with 404:
llm.example.com {
handle /v1/* {
reverse_proxy 127.0.0.1:8000
}
handle {
respond 404
}
}sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enableClients now use https://llm.example.com/v1 as the base URL and the value of VLLM_API_KEY as their API key. Streaming responses pass through Caddy without extra settings.
Alternative: install vLLM with uv and run it under systemd
If you cannot use Docker, the documentation recommends installing vLLM with uv in a fresh virtual environment. Install uv with its official installer after reviewing it, then create the environment in /opt/vllm with a managed Python 3.12 stored inside the project directory, so the service user can read it:
sudo apt install build-essential
curl -LsSf https://astral.sh/uv/install.sh -o uv-install.sh
less uv-install.sh
sh uv-install.sh
source $HOME/.local/bin/env
sudo mkdir -p /opt/vllm && sudo chown $USER:$USER /opt/vllm
cd /opt/vllm
export UV_PYTHON_INSTALL_DIR=/opt/vllm/python
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
python -c "import vllm; print(vllm.__version__)"--torch-backend=auto selects the PyTorch build that matches your installed driver. Next, create a service user and a root-only environment file:
sudo useradd --system --create-home --home-dir /var/lib/vllm --shell /usr/sbin/nologin vllm
echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/vllm.env > /dev/null
echo "HF_HOME=/var/lib/vllm/huggingface" | sudo tee -a /etc/vllm.env > /dev/null
sudo chmod 600 /etc/vllm.envCreate /etc/systemd/system/vllm.service, then run sudo systemctl daemon-reload and sudo systemctl enable --now vllm:
[Unit]
Description=vLLM OpenAI-compatible server
After=network-online.target
Wants=network-online.target
[Service]
User=vllm
Group=vllm
EnvironmentFile=/etc/vllm.env
ExecStart=/opt/vllm/.venv/bin/vllm serve Qwen/Qwen3-0.6B --host 127.0.0.1 --port 8000 --max-model-len 8192
Restart=on-failure
RestartSec=10
[Install]
WantedBy=multi-user.targetThe --host 127.0.0.1 option keeps the server on loopback. Follow the start-up with journalctl -u vllm -f, test it as in Step 5 and use the same Caddy block. Add HF_TOKEN= to /etc/vllm.env for gated models.
Back up and restore
vLLM keeps no user data. Back up the configuration: compose.yaml and .env (or /etc/vllm.env and the unit file). The Hugging Face cache can be downloaded again; include it only if downloads are slow or the model is private.
sudo mkdir -p /opt/backups
sudo tar czf /opt/backups/vllm-config-$(date +%F).tar.gz /opt/vllm/compose.yaml /opt/vllm/.envTo restore, prepare a server with Steps 1 and 2, extract the archive with sudo tar xzf /opt/backups/vllm-config-YYYY-MM-DD.tar.gz -C /, fix the ownership of /opt/vllm and run docker compose up -d. Copy the archive off the server; it contains your API key.
Update vLLM
Read the release notes on the vLLM releases page first: options and defaults change between versions. For Docker, change the pinned tag (or keep latest), pull and recreate:
cd /opt/vllm
docker compose pull
docker compose up -d
docker image pruneFor the uv installation, the documentation recommends a fresh environment rather than upgrading in place, because compiled kernels are tied to specific CUDA and PyTorch versions. Create a new virtual environment next to the old one, install vLLM into it, point ExecStart at the new path and restart the service; keep the old environment until the new one works. Driver updates arrive through apt upgrade from NVIDIA's repository and need a reboot.
Troubleshooting
CUDA out of memory
The weights plus KV cache do not fit. Lower --max-model-len or --max-num-seqs, use a quantized version of the model, split it with --tensor-parallel-size, or check with nvidia-smi that no other process holds GPU memory. If other workloads share the GPU, lower --gpu-memory-utilization. The vLLM memory guide also suggests --enforce-eager to skip CUDA graph memory.
could not select device driver with capabilities gpu
Docker does not know the NVIDIA runtime. Repeat Step 2, especially sudo nvidia-ctk runtime configure --runtime=docker and the Docker restart, and test with the sample workload.
CUDA driver is too old or a PTX toolchain error
The driver is older than the CUDA version of the image or wheel. Update the driver from NVIDIA's repository and reboot. For some datacenter GPUs the vLLM documentation offers a compatibility mode: add VLLM_ENABLE_CUDA_COMPATIBILITY: "1" to the container environment.
401 or 403 when downloading a gated model
HF_TOKEN is missing or invalid, or your Hugging Face account has not been granted access to the model yet. Fix the token in .env, request access on the model's Hugging Face page, wait until it is granted and run docker compose up -d.
The server hangs while downloading or starting
Download the model separately with the Hugging Face hf command line tool into the cache directory and start vLLM again; this shows whether the download is the problem. For more output, set VLLM_LOGGING_LEVEL: "DEBUG" in the environment, and remove it again when you are done.
Next steps
- Add a chat interface by connecting Open WebUI to
https://llm.example.com/v1. - Compare with CPU-friendly runtimes: llama.cpp server and Ollama.
- Choose hardware on the GPU servers and LLM API hosting pages.
- Read the official documentation at https://docs.vllm.ai for every engine argument.
Frequently asked questions
Which GPUs does vLLM support?
vLLM's CUDA builds need an NVIDIA GPU with compute capability 7.5 or higher, for example T4, RTX 20 series and newer, A100, L4, H100 or B200. It runs on Linux only; there is no native Windows support.
Is the vLLM API key enough to secure the server?
No. The vLLM documentation states that --api-key only authenticates endpoints under /v1, /v2, /inference and /cohere, while other endpoints on the same server stay open. Keep vLLM on 127.0.0.1 and let a reverse proxy expose only the paths you need, as this guide does with Caddy.
How much GPU memory does a model need?
The model weights must fit in GPU memory, and vLLM uses the rest of its share for the KV cache. By default an instance may use 92 percent of each GPU’s memory (--gpu-memory-utilization 0.92). Larger models need quantized weights or tensor parallelism across several GPUs.
How do I serve a gated model such as Llama?
Request access on the model’s Hugging Face page with your account, create an access token and put it in HF_TOKEN in the .env file. Once access is granted, vLLM downloads the weights into the Hugging Face cache volume.
Should I use Docker or pip to install vLLM?
The official vllm/vllm-openai image bundles a matching CUDA and PyTorch stack and is the simplest way to run the server. The pip or uv install in a virtual environment suits custom setups; the docs recommend a fresh environment because compiled kernels are tied to specific CUDA and PyTorch versions.
Sources
- docs.vllm.ai/en/latest/getting_started/installation/gpu
- docs.vllm.ai/en/latest/deployment/docker
- docs.vllm.ai/en/latest/serving/online_serving/openai_compatible_server
- docs.vllm.ai/en/latest/serving/online_serving
- docs.vllm.ai/en/latest/cli/serve
- docs.vllm.ai/en/latest/configuration/engine_args
- docs.vllm.ai/en/latest/configuration/conserving_memory
- docs.vllm.ai/en/latest/usage/security
- docs.vllm.ai/en/latest/usage/troubleshooting
- docs.nvidia.com/datacenter/tesla/driver-installation-guide/ubuntu.html
- docs.nvidia.com/datacenter/tesla/driver-installation-guide/debian.html
- tracker.debian.org/pkg/software-properties