# How to run llama.cpp server as an OpenAI-compatible API

> Build llama.cpp from source for CPU or NVIDIA CUDA, serve GGUF models with llama-server and an API key, run it under systemd and put Caddy in front for HTTPS.

Difficulty: Intermediate\
Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13

llama.cpp is an open-source C/C++ engine for running large language models in the GGUF format on CPUs and GPUs. Its `llama-server` program provides an OpenAI-compatible HTTP API (`/v1/chat/completions`, `/v1/completions`, `/v1/embeddings`) and a small built-in web interface. This guide **builds llama.cpp from source** with CMake, for CPU or with NVIDIA CUDA, tests `llama-server` with a model downloaded from Hugging Face, runs it as a **systemd service under a dedicated user** on `127.0.0.1` with an API key, and publishes it over HTTPS with **Caddy**. A Docker alternative with the official container images is included.

> **Note**
>
> The llama.cpp README also offers pre-built binaries on the GitHub releases page and an install script. Building from source, as shown here, lets you choose the CUDA options and update on your own schedule.

## Prerequisites

- A server running **Ubuntu 24.04 LTS**, **Ubuntu 26.04 LTS**, **Debian 12** or **Debian 13**.
- A non-root user with `sudo` rights; see [Secure a new Linux server](/guides/secure-a-new-linux-server) and [Set up SSH keys](/guides/ssh-keys).
- Optional, for GPU offload: an NVIDIA GPU with a working driver (`nvidia-smi` prints your GPU). Step 1 of [How to install vLLM](/guides/install-vllm) shows how to install the driver from NVIDIA's repository; see [GPU servers](/gpu-servers) for suitable hardware.
- A domain such as `llm.example.com` pointing at the server and Caddy installed from [Caddy as a reverse proxy](/guides/caddy-reverse-proxy), if clients outside the server need the API.

| Resource | Minimum (official) | Suggested starting point |
|---|---|---|
| CPU | Not published | 4 vCPU; generation uses the physical cores |
| RAM | Not published | At least the GGUF file size plus 2 to 4 GB for context and the system |
| GPU (optional) | NVIDIA GPU with the CUDA toolkit installed for the CUDA build | Enough VRAM to hold the model so all layers run on the GPU |
| Disk | Not published | 20 GB for the source, build and a few small models |

The llama.cpp project does not publish minimum requirements. The suggested values are a conservative starting point based on the rule of thumb that the model file has to fit in memory, not a benchmark. Larger context sizes and more parallel slots need more memory.

## Step 1 — Install the build tools

llama.cpp needs a C/C++ compiler, CMake and Git. Install the OpenSSL development package too: llama.cpp's HTTPS support, which `-hf` downloads from Hugging Face use, is built with OpenSSL and is enabled by default when the library is present.

```bash
sudo apt update
sudo apt install build-essential cmake git libssl-dev
```

Clone the repository into `/opt/llama.cpp`, owned by your admin user:

```bash
sudo mkdir -p /opt/llama.cpp && sudo chown $USER:$USER /opt/llama.cpp
git clone https://github.com/ggml-org/llama.cpp /opt/llama.cpp
cd /opt/llama.cpp
```

## Step 2 — Build llama.cpp

Pick the CPU build, or the CUDA build if the server has an NVIDIA GPU. The CUDA build needs the CUDA toolkit, which provides the `nvcc` compiler; it comes from the same NVIDIA repository as the driver.

**CPU**

```bash
cd /opt/llama.cpp
cmake -B build
cmake --build build --config Release -j $(nproc)
```
**NVIDIA CUDA**

```bash
sudo apt install cuda-toolkit
export PATH=/usr/local/cuda/bin:$PATH
nvcc --version
cd /opt/llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j $(nproc)
```

The build takes several minutes. For CUDA builds, the build documentation lets you limit compilation to your GPU's compute capability, for example `-DCMAKE_CUDA_ARCHITECTURES="86;89"`, which shortens the build. NVIDIA's installation guide places the toolkit in a versioned directory such as `/usr/local/cuda-13.4`; `/usr/local/cuda` normally points to the installed version. Binaries end up in `build/bin`. Check that the server program runs:

```bash
./build/bin/llama-server --version
```

## Step 3 — Run a first test with a GGUF model

`llama-server` loads a model either from a local GGUF file with `-m` or straight from Hugging Face with `-hf user/model[:quant]`; without a quant tag it picks `Q4_K_M`, or the first file in the repository. The project's README uses the small `ggml-org/Qwen3.5-0.8B-GGUF` model as an example, which is ideal for a first test:

```bash
cd /opt/llama.cpp
./build/bin/llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF --host 127.0.0.1 --port 8080 -c 8192
```

The server downloads the model into your Hugging Face cache (`~/.cache/huggingface/hub`), loads it and prints that it is listening on `127.0.0.1:8080`. In a second SSH session, test the endpoints:

```bash
curl http://127.0.0.1:8080/health
curl http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Write a haiku about servers"}]}'
```

`/health` returns `{"status":"ok"}` once the model is loaded (503 while loading), and the chat request returns a JSON completion. Stop the test server with `Ctrl+C`. The most important options:

| Option | Default | Purpose |
|---|---|---|
| `-m`, `--model` | none | Path to a local GGUF file |
| `-hf`, `--hf-repo` | none | Download a model from Hugging Face as `user/model[:quant]` |
| `-c`, `--ctx-size` | `0` (taken from the model) | Context size in tokens; set it explicitly to cap memory use |
| `-t`, `--threads` | `-1` (automatic) | CPU threads used for generation |
| `-ngl`, `--n-gpu-layers` | `auto` | Layers stored in VRAM: a number, `auto` or `all` |
| `-np`, `--parallel` | `-1` (automatic) | Number of server slots for parallel requests |
| `--host`, `--port` | `127.0.0.1`, `8080` | Listen address and port |
| `--api-key` | none | API key or comma-separated keys; also `LLAMA_API_KEY` |
| `--no-webui` | web UI on | Disables the built-in web interface |

Every option also has an environment variable (for example `LLAMA_ARG_CTX_SIZE`), listed in the server README.

## Step 4 — Create a service user and the API key

Run the server under its own unprivileged user with its own model directory, and keep the API key in a root-only environment file so it does not show up in the process list:

```bash
sudo useradd --system --create-home --home-dir /var/lib/llama --shell /usr/sbin/nologin llama
sudo install -d -o llama -g llama /var/lib/llama/models
echo "LLAMA_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/llama-server.env > /dev/null
echo "LLAMA_CACHE=/var/lib/llama/models" | sudo tee -a /etc/llama-server.env > /dev/null
sudo chmod 600 /etc/llama-server.env
sudo cat /etc/llama-server.env
```

`LLAMA_CACHE` tells llama.cpp where to store `-hf` downloads, so the service keeps its models in `/var/lib/llama/models`. Note the API key; clients need it. The service downloads the model again into its own directory, so you can remove the test copy from `~/.cache/huggingface/hub` afterwards.

## Step 5 — Run llama-server with systemd

Create `/etc/systemd/system/llama-server.service`:

```ini
[Unit]
Description=llama.cpp server
After=network-online.target
Wants=network-online.target

[Service]
User=llama
Group=llama
EnvironmentFile=/etc/llama-server.env
ExecStart=/opt/llama.cpp/build/bin/llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF --host 127.0.0.1 --port 8080 -c 8192
Restart=on-failure
RestartSec=5

[Install]
WantedBy=multi-user.target
```

To serve a GGUF file you copied to the server yourself, put it in `/var/lib/llama/models`, make the `llama` user its owner and replace `-hf ggml-org/Qwen3.5-0.8B-GGUF` with `-m /var/lib/llama/models/your-model.gguf`. Enable and start the service, and follow the log until the model is loaded:

```bash
sudo systemctl daemon-reload
sudo systemctl enable --now llama-server
sudo journalctl -u llama-server -f
```

Now check that the API key is enforced. The first request must fail with `401`, the second must list the model:

```bash
KEY=$(sudo grep LLAMA_API_KEY /etc/llama-server.env | cut -d= -f2)
curl -i http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/models -H "Authorization: Bearer $KEY"
```

With a key configured, the API endpoints require it; `/health` stays public so monitoring can check the server without credentials.

## Step 6 — Publish the API over HTTPS with Caddy

Add a site block to `/etc/caddy/Caddyfile`, reload Caddy and allow only SSH and web traffic in the firewall:

```caddyfile
llm.example.com {
    reverse_proxy 127.0.0.1:8080
}
```

```bash
sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
```

OpenAI-compatible clients now use `https://llm.example.com/v1` as the base URL and your key as the API key. Caddy streams the server-sent events of streaming responses without extra settings. If you prefer Nginx, follow [Nginx with Certbot](/guides/nginx-reverse-proxy-certbot) and turn off proxy buffering for the API location.

> **Warning**
>
> Never start `llama-server` with `--host 0.0.0.0` on a public server without an API key and a firewall. Anyone who reaches the port can use your CPU or GPU for their own requests.

## Run llama-server with Docker instead

The project publishes container images on GitHub Container Registry: `ghcr.io/ggml-org/llama.cpp:server` for CPU, `:server-cuda` (CUDA 12) and `:server-cuda13` for NVIDIA GPUs, plus ROCm, Vulkan and Intel variants. The server image already listens on all container addresses, so publish the port on loopback only. Create `/opt/llama-server/compose.yaml` and an `.env` file with `LLAMA_API_KEY=` and a key from `openssl rand -hex 32` (`chmod 600 .env`):

```yaml
services:
  llama-server:
    image: ghcr.io/ggml-org/llama.cpp:server
    command: ["-hf", "ggml-org/Qwen3.5-0.8B-GGUF", "--port", "8080", "-c", "8192"]
    environment:
      LLAMA_API_KEY: "${LLAMA_API_KEY}"
      LLAMA_CACHE: "/models"
    volumes:
      - ./models:/models
    ports:
      - "127.0.0.1:8080:8080"
    restart: unless-stopped
```

```bash
cd /opt/llama-server
docker compose up -d
docker compose logs -f
```

For NVIDIA GPUs, use `ghcr.io/ggml-org/llama.cpp:server-cuda`, install the NVIDIA Container Toolkit (Step 2 of [How to install vLLM](/guides/install-vllm)) and add a `deploy.resources.reservations.devices` block with `driver: nvidia`, `count: all` and `capabilities: [gpu]` to the service. The documentation notes that the GPU images are built by CI and not tested beyond that. Use either Docker or the systemd service on port 8080, not both.

## Back up and restore

The build itself can be recreated at any time. What you need to keep is the configuration and, if downloads are slow or the files are your own, the models:

```bash
sudo mkdir -p /opt/backups
sudo tar czf /opt/backups/llama-server-$(date +%F).tar.gz /etc/llama-server.env /etc/systemd/system/llama-server.service /var/lib/llama/models
```

The archive contains your API key, so store it as carefully as the key itself. To restore on a new server, build llama.cpp as in Steps 1 and 2, create the `llama` user from Step 4, then extract the archive and restart the service:

```bash
sudo tar xzf /opt/backups/llama-server-YYYY-MM-DD.tar.gz -C /
sudo chown -R llama:llama /var/lib/llama
sudo systemctl daemon-reload
sudo systemctl enable --now llama-server
```

Copy the archive off the server as well.

## Update llama.cpp

llama.cpp changes quickly and publishes frequent numbered builds on its GitHub releases page; read the release notes for changes to server options before you update. Pull the latest code, rebuild with the same options (CMake keeps them in the `build` directory) and restart:

```bash
cd /opt/llama.cpp
git pull
cmake -B build
cmake --build build --config Release -j $(nproc)
sudo systemctl restart llama-server
```

If a new version misbehaves, check out the previous release tag with `git checkout` and rebuild. For the Docker variant, run `docker compose pull` and `docker compose up -d`.

## Troubleshooting

### CMake cannot find CUDA or nvcc is not found

The CUDA toolkit is missing or not in your `PATH`. Install `cuda-toolkit` from NVIDIA's repository, run `export PATH=/usr/local/cuda/bin:$PATH` and check `nvcc --version`. Then delete the `build` directory and configure again with `-DGGML_CUDA=ON`.

### 401 Unauthorized from the API

The server has an API key and the request did not send it, or sent a different one. Send `Authorization: Bearer` followed by the key from `/etc/llama-server.env`. In OpenAI client libraries, set the key as `api_key`.

### Downloads with -hf fail

llama.cpp was built without HTTPS support, usually because `libssl-dev` was not installed. Install it, delete the `build` directory and build again. For gated or private repositories, add `HF_TOKEN=` with a Hugging Face access token to `/etc/llama-server.env`.

### Out of memory or failed to allocate when loading a model

The model and its context do not fit in RAM or VRAM. Lower `-c`, choose a smaller quantization (for example `:Q4_K_M` instead of `:Q8_0`), use a smaller model, reduce `-np`, or set `-ngl` to a number below the model's layer count so part of the model stays in RAM.

### The service fails with permission denied

The `llama` user cannot read the binary or write to its cache. Check with `sudo -u llama ls /opt/llama.cpp/build/bin` and `sudo -u llama touch /var/lib/llama/models/test`, and fix ownership with `sudo chown -R llama:llama /var/lib/llama`.

## Next steps

- Connect a chat interface: [Open WebUI](/guides/open-webui-ollama) can add `https://llm.example.com/v1` as an OpenAI-compatible connection.
- Compare with the managed approach in [How to install Ollama](/guides/install-ollama).
- For high-throughput serving on NVIDIA GPUs, see [How to install vLLM](/guides/install-vllm).
- Compare servers for your own API on the [LLM API hosting](/llm-api-hosting) page.
- Read the server documentation at https://github.com/ggml-org/llama.cpp/tree/master/tools/server for every option.

## Frequently asked questions

### What is the difference between llama.cpp and Ollama?

llama.cpp is the inference engine and llama-server is its HTTP server. You choose the GGUF file, context size, threads and GPU layers yourself. Ollama adds a model library and its own management layer on top. Use llama-server when you want direct control over these settings.

### Do I need a GPU for llama-server?

No. llama.cpp runs GGUF models on the CPU. A CUDA build offloads layers to an NVIDIA GPU, and with the default -ngl auto it places as many layers in VRAM as fit.

### Is the llama-server API protected?

Only if you set an API key with --api-key, --api-key-file or the LLAMA_API_KEY variable. With a key set, API requests need an Authorization: Bearer header, while the health endpoint stays public for monitoring. Keep the server on 127.0.0.1 behind a reverse proxy in any case.

### Where are models downloaded with -hf stored?

In the Hugging Face cache of the user running the server, by default ~/.cache/huggingface/hub. Set LLAMA_CACHE to use a different directory, as this guide does for the service user.

### How large a model can my server run?

The llama.cpp project does not publish requirements. As a rule of thumb, you need at least the size of the GGUF file in RAM or VRAM plus memory for the context, so start with a small quantized model and a fixed context size.

---

Source: <https://hyperdc.com/guides/tutorials/llama-cpp-server>\
Updated: 2026-10-09
