# How to install vLLM on an NVIDIA GPU server with Docker

> Install the NVIDIA driver and Container Toolkit, run the vLLM OpenAI-compatible server with Docker Compose, secure it with an API key and Caddy, tune memory.

Difficulty: Advanced\
Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13

vLLM is a high-throughput inference and serving engine for large language models. It keeps many requests in flight on the GPU at once and exposes an OpenAI-compatible HTTP server, which makes it a common choice for production LLM APIs on NVIDIA hardware. This guide prepares a GPU server with **NVIDIA's driver and Container Toolkit**, runs the official **`vllm/vllm-openai` image with Docker Compose** on `127.0.0.1:8000`, protects it with an API key and a **Caddy** proxy that exposes only the `/v1` API, explains the memory and multi-GPU options, and shows an alternative installation with `uv` and systemd.

## Prerequisites

- A dedicated server with a supported NVIDIA GPU; see [GPU servers](/gpu-servers). vLLM's CUDA builds need compute capability 7.5 or newer.
- **Ubuntu 24.04 LTS**, **Ubuntu 26.04 LTS**, **Debian 12** or **Debian 13** on x86_64. NVIDIA's driver guide covers all four releases.
- A non-root user with `sudo` rights; see [Secure a new Linux server](/guides/secure-a-new-linux-server) and [Set up SSH keys](/guides/ssh-keys).
- Docker Engine with the Compose plugin from [Install Docker on Ubuntu](/guides/install-docker-ubuntu) or [Install Docker on Debian](/guides/install-docker-debian).
- A Hugging Face account and access token if you want to serve gated models.
- A domain such as `llm.example.com` pointing at the server, and Caddy from [Caddy as a reverse proxy](/guides/caddy-reverse-proxy).

| Resource | Minimum (official) | Suggested starting point |
|---|---|---|
| GPU | NVIDIA, compute capability 7.5 or higher | Enough VRAM for the model weights plus KV cache |
| Driver and CUDA | Default builds use CUDA 12.9; CUDA 13 images need driver R580 or newer | Current driver from NVIDIA's repository |
| Python (pip method) | 3.11 to 3.14 | 3.12, as in the official uv example |
| RAM | Not published | At least as much system RAM as GPU memory |
| Disk | Not published | 100 GB free for the image and model weights |

The vLLM documentation does not publish RAM or disk minimums; the suggested values are a conservative starting point, not a benchmark. Check the size of a model's weight files on its Hugging Face page before you choose a GPU.

## Step 1 — Install the NVIDIA driver

Install the driver from NVIDIA's network repository as described in NVIDIA's driver installation guide. The commands add the `cuda-keyring` package for your release and install the `nvidia-open` driver. On Debian, NVIDIA's guide also requires the `contrib` component. The guide enables it with `add-apt-repository`, which Debian 13 no longer ships, so first add `contrib` next to `main` in your APT sources yourself (the `Components:` line in `/etc/apt/sources.list.d/debian.sources`, or the `deb` lines in `/etc/apt/sources.list`, depending on which file your system uses):

**Ubuntu**

```bash
sudo apt update
sudo apt install linux-headers-$(uname -r)
distro=ubuntu$(. /etc/os-release && echo "$VERSION_ID" | tr -d .)
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install nvidia-open
sudo reboot
```
**Debian**

```bash
sudo apt update
sudo apt install linux-headers-$(uname -r)
distro=debian$(. /etc/os-release && echo "$VERSION_ID")
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt -V install nvidia-open
sudo reboot
```

After the reboot, run `nvidia-smi`. It lists every GPU, the driver version and, in the header, the highest CUDA version the driver supports. If your GPU needs an older driver branch, NVIDIA's guide explains how to pin one. If Secure Boot is enabled, the kernel module must be signed with a key the firmware trusts before it can load.

## Step 2 — Install the NVIDIA Container Toolkit

The Container Toolkit lets Docker containers use the GPU. Add NVIDIA's repository, install the toolkit and configure Docker's runtime, following NVIDIA's installation guide:

```bash
sudo apt install -y --no-install-recommends ca-certificates curl gnupg2
curl -fsSL https://nvidia.github.io/libnvidia-container/gpgkey | sudo gpg --dearmor -o /usr/share/keyrings/nvidia-container-toolkit-keyring.gpg
curl -s -L https://nvidia.github.io/libnvidia-container/stable/deb/nvidia-container-toolkit.list | sed 's#deb https://#deb [signed-by=/usr/share/keyrings/nvidia-container-toolkit-keyring.gpg] https://#g' | sudo tee /etc/apt/sources.list.d/nvidia-container-toolkit.list
sudo apt update
sudo apt install -y nvidia-container-toolkit
sudo nvidia-ctk runtime configure --runtime=docker
sudo systemctl restart docker
```

`nvidia-ctk` updates `/etc/docker/daemon.json` so that Docker can use the NVIDIA Container Runtime. Run NVIDIA's sample workload to confirm that containers see the GPU:

```bash
sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi
```

You should see the same `nvidia-smi` table as on the host.

## Step 3 — Create the project directory and secrets

Keep the Compose file, secrets and the Hugging Face cache together in `/opt/vllm`:

```bash
sudo mkdir -p /opt/vllm/hf-cache && sudo chown -R $USER:$USER /opt/vllm
cd /opt/vllm
echo "VLLM_API_KEY=$(openssl rand -hex 32)" > .env
echo "HF_TOKEN=" >> .env
chmod 600 .env
```

`VLLM_API_KEY` is read by the vLLM server as its API key. Leave `HF_TOKEN` empty for public models, or paste a Hugging Face access token after the `=` for gated models; for those, also request access on the model's Hugging Face page.

## Step 4 — Start vLLM with Docker Compose

Create `/opt/vllm/compose.yaml`. The service follows the documented `docker run` command: all GPUs, the host IPC namespace (PyTorch shares memory between processes, which needs `--ipc=host` or a larger `--shm-size`), the Hugging Face cache mounted into the container and the API on port 8000, published on loopback only. This example serves the small `Qwen/Qwen3-0.6B` model used in the vLLM documentation:

```yaml
services:
  vllm:
    image: vllm/vllm-openai:latest
    command: ["--model", "Qwen/Qwen3-0.6B", "--max-model-len", "8192"]
    environment:
      VLLM_API_KEY: "${VLLM_API_KEY}"
      HF_TOKEN: "${HF_TOKEN}"
    volumes:
      - ./hf-cache:/root/.cache/huggingface
    ports:
      - "127.0.0.1:8000:8000"
    ipc: host
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
    restart: unless-stopped
```

```bash
docker compose up -d
docker compose logs -f vllm
```

The first start pulls a large image, downloads the model weights into `/opt/vllm/hf-cache` and prepares the GPU kernels, which can take several minutes. The server is ready when the log shows that the application started and Uvicorn is serving on port 8000.

> **Tip**
>
> `latest` follows each new release. For production, pin a release tag such as `vllm/vllm-openai:vX.Y.Z` from the vLLM releases page on GitHub, and change it deliberately when you update.

## Step 5 — Test the OpenAI-compatible API

Load the key into your shell and query the server. `/health` needs no key; `/v1` endpoints do:

```bash
cd /opt/vllm
export VLLM_API_KEY=$(grep VLLM_API_KEY .env | cut -d= -f2)
curl -i http://127.0.0.1:8000/health
curl http://127.0.0.1:8000/v1/models -H "Authorization: Bearer $VLLM_API_KEY"
curl http://127.0.0.1:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer $VLLM_API_KEY" \
  -d '{"model": "Qwen/Qwen3-0.6B", "messages": [{"role": "user", "content": "Say hello in five words"}]}'
```

`/health` returns `200 OK`, `/v1/models` lists `Qwen/Qwen3-0.6B`, and the chat request returns a completion. A request to `/v1/models` without the header returns `401`. Use `--served-model-name` if clients expect a different model name.

## Step 6 — Tune memory and use several GPUs

vLLM reserves a share of each GPU's memory when it starts, loads the weights into it and uses the rest as KV cache for concurrent requests. Add these options to `command` in `compose.yaml` and run `docker compose up -d` to apply them:

| Option | Default | Use it to |
|---|---|---|
| `--gpu-memory-utilization` | `0.92` | Set the fraction of GPU memory this instance may use; lower it if other processes share the GPU |
| `--max-model-len` | From the model config | Cap the context length (prompt plus output); `auto` picks the largest length that fits |
| `--max-num-seqs` | Set by vLLM | Limit how many sequences are processed per iteration, which reduces memory use |
| `--tensor-parallel-size`, `-tp` | `1` | Split a model across several GPUs in one server |
| `--quantization`, `-q` | From the model config | Select a quantization method for smaller weights |
| `--enforce-eager` | off | Disable CUDA graphs, which saves some GPU memory |

For example, to spread a larger model across two GPUs, use `command: ["--model", "org/model-name", "--tensor-parallel-size", "2"]`. To dedicate specific GPUs to one instance, set `CUDA_VISIBLE_DEVICES` in `environment`, for example `CUDA_VISIBLE_DEVICES: "0,1"`. Multi-node deployments need an isolated network between the nodes, because vLLM's internal communication is not secured.

## Step 7 — Expose only the API through Caddy

The vLLM security documentation recommends a reverse proxy that allows only the endpoints you want to expose, because `--api-key` protects `/v1`, `/v2`, `/inference` and `/cohere` but not endpoints such as `/tokenize`, `/pooling` or `/health`. This Caddy site block forwards `/v1/*` and answers everything else with 404:

```caddyfile
llm.example.com {
    handle /v1/* {
        reverse_proxy 127.0.0.1:8000
    }
    handle {
        respond 404
    }
}
```

```bash
sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
```

Clients now use `https://llm.example.com/v1` as the base URL and the value of `VLLM_API_KEY` as their API key. Streaming responses pass through Caddy without extra settings.

> **Warning**
>
> Do not publish port 8000 on all addresses (`"8000:8000"`). Docker-published ports bypass ufw, and the unauthenticated endpoints would be reachable from the internet.

## Alternative: install vLLM with uv and run it under systemd

If you cannot use Docker, the documentation recommends installing vLLM with `uv` in a fresh virtual environment. Install `uv` with its official installer after reviewing it, then create the environment in `/opt/vllm` with a managed Python 3.12 stored inside the project directory, so the service user can read it:

```bash
sudo apt install build-essential
curl -LsSf https://astral.sh/uv/install.sh -o uv-install.sh
less uv-install.sh
sh uv-install.sh
source $HOME/.local/bin/env
sudo mkdir -p /opt/vllm && sudo chown $USER:$USER /opt/vllm
cd /opt/vllm
export UV_PYTHON_INSTALL_DIR=/opt/vllm/python
uv venv --python 3.12 --seed --managed-python
source .venv/bin/activate
uv pip install vllm --torch-backend=auto
python -c "import vllm; print(vllm.__version__)"
```

`--torch-backend=auto` selects the PyTorch build that matches your installed driver. Next, create a service user and a root-only environment file:

```bash
sudo useradd --system --create-home --home-dir /var/lib/vllm --shell /usr/sbin/nologin vllm
echo "VLLM_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/vllm.env > /dev/null
echo "HF_HOME=/var/lib/vllm/huggingface" | sudo tee -a /etc/vllm.env > /dev/null
sudo chmod 600 /etc/vllm.env
```

Create `/etc/systemd/system/vllm.service`, then run `sudo systemctl daemon-reload` and `sudo systemctl enable --now vllm`:

```ini
[Unit]
Description=vLLM OpenAI-compatible server
After=network-online.target
Wants=network-online.target

[Service]
User=vllm
Group=vllm
EnvironmentFile=/etc/vllm.env
ExecStart=/opt/vllm/.venv/bin/vllm serve Qwen/Qwen3-0.6B --host 127.0.0.1 --port 8000 --max-model-len 8192
Restart=on-failure
RestartSec=10

[Install]
WantedBy=multi-user.target
```

The `--host 127.0.0.1` option keeps the server on loopback. Follow the start-up with `journalctl -u vllm -f`, test it as in Step 5 and use the same Caddy block. Add `HF_TOKEN=` to `/etc/vllm.env` for gated models.

## Back up and restore

vLLM keeps no user data. Back up the configuration: `compose.yaml` and `.env` (or `/etc/vllm.env` and the unit file). The Hugging Face cache can be downloaded again; include it only if downloads are slow or the model is private.

```bash
sudo mkdir -p /opt/backups
sudo tar czf /opt/backups/vllm-config-$(date +%F).tar.gz /opt/vllm/compose.yaml /opt/vllm/.env
```

To restore, prepare a server with Steps 1 and 2, extract the archive with `sudo tar xzf /opt/backups/vllm-config-YYYY-MM-DD.tar.gz -C /`, fix the ownership of `/opt/vllm` and run `docker compose up -d`. Copy the archive off the server; it contains your API key.

## Update vLLM

Read the release notes on the vLLM releases page first: options and defaults change between versions. For Docker, change the pinned tag (or keep `latest`), pull and recreate:

```bash
cd /opt/vllm
docker compose pull
docker compose up -d
docker image prune
```

For the uv installation, the documentation recommends a fresh environment rather than upgrading in place, because compiled kernels are tied to specific CUDA and PyTorch versions. Create a new virtual environment next to the old one, install vLLM into it, point `ExecStart` at the new path and restart the service; keep the old environment until the new one works. Driver updates arrive through `apt upgrade` from NVIDIA's repository and need a reboot.

## Troubleshooting

### CUDA out of memory

The weights plus KV cache do not fit. Lower `--max-model-len` or `--max-num-seqs`, use a quantized version of the model, split it with `--tensor-parallel-size`, or check with `nvidia-smi` that no other process holds GPU memory. If other workloads share the GPU, lower `--gpu-memory-utilization`. The vLLM memory guide also suggests `--enforce-eager` to skip CUDA graph memory.

### could not select device driver with capabilities gpu

Docker does not know the NVIDIA runtime. Repeat Step 2, especially `sudo nvidia-ctk runtime configure --runtime=docker` and the Docker restart, and test with the sample workload.

### CUDA driver is too old or a PTX toolchain error

The driver is older than the CUDA version of the image or wheel. Update the driver from NVIDIA's repository and reboot. For some datacenter GPUs the vLLM documentation offers a compatibility mode: add `VLLM_ENABLE_CUDA_COMPATIBILITY: "1"` to the container environment.

### 401 or 403 when downloading a gated model

`HF_TOKEN` is missing or invalid, or your Hugging Face account has not been granted access to the model yet. Fix the token in `.env`, request access on the model's Hugging Face page, wait until it is granted and run `docker compose up -d`.

### The server hangs while downloading or starting

Download the model separately with the Hugging Face `hf` command line tool into the cache directory and start vLLM again; this shows whether the download is the problem. For more output, set `VLLM_LOGGING_LEVEL: "DEBUG"` in the environment, and remove it again when you are done.

## Next steps

- Add a chat interface by connecting [Open WebUI](/guides/open-webui-ollama) to `https://llm.example.com/v1`.
- Compare with CPU-friendly runtimes: [llama.cpp server](/guides/llama-cpp-server) and [Ollama](/guides/install-ollama).
- Choose hardware on the [GPU servers](/gpu-servers) and [LLM API hosting](/llm-api-hosting) pages.
- Read the official documentation at https://docs.vllm.ai for every engine argument.

## Frequently asked questions

### Which GPUs does vLLM support?

vLLM's CUDA builds need an NVIDIA GPU with compute capability 7.5 or higher, for example T4, RTX 20 series and newer, A100, L4, H100 or B200. It runs on Linux only; there is no native Windows support.

### Is the vLLM API key enough to secure the server?

No. The vLLM documentation states that --api-key only authenticates endpoints under /v1, /v2, /inference and /cohere, while other endpoints on the same server stay open. Keep vLLM on 127.0.0.1 and let a reverse proxy expose only the paths you need, as this guide does with Caddy.

### How much GPU memory does a model need?

The model weights must fit in GPU memory, and vLLM uses the rest of its share for the KV cache. By default an instance may use 92 percent of each GPU’s memory (--gpu-memory-utilization 0.92). Larger models need quantized weights or tensor parallelism across several GPUs.

### How do I serve a gated model such as Llama?

Request access on the model’s Hugging Face page with your account, create an access token and put it in HF_TOKEN in the .env file. Once access is granted, vLLM downloads the weights into the Hugging Face cache volume.

### Should I use Docker or pip to install vLLM?

The official vllm/vllm-openai image bundles a matching CUDA and PyTorch stack and is the simplest way to run the server. The pip or uv install in a virtual environment suits custom setups; the docs recommend a fresh environment because compiled kernels are tied to specific CUDA and PyTorch versions.

---

Source: <https://hyperdc.com/guides/tutorials/install-vllm>\
Updated: 2026-10-09
