Skip to content

TutorialsAI & LLM

How to run llama.cpp server as an OpenAI-compatible API

Build llama.cpp from source for CPU or NVIDIA CUDA, serve GGUF models with llama-server and an API key, run it under systemd and put Caddy in front for HTTPS.

  • Intermediate
  • 40 min read
  • Updated

Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13

This guide is not available in your language yet, so it is shown in English.

On this page
  1. Prerequisites
  2. Step 1 — Install the build tools
  3. Step 2 — Build llama.cpp
  4. Step 3 — Run a first test with a GGUF model
  5. Step 4 — Create a service user and the API key
  6. Step 5 — Run llama-server with systemd
  7. Step 6 — Publish the API over HTTPS with Caddy
  8. Run llama-server with Docker instead
  9. Back up and restore
  10. Update llama.cpp
  11. Troubleshooting
  12. CMake cannot find CUDA or nvcc is not found
  13. 401 Unauthorized from the API
  14. Downloads with -hf fail
  15. Out of memory or failed to allocate when loading a model
  16. The service fails with permission denied
  17. Next steps

llama.cpp is an open-source C/C++ engine for running large language models in the GGUF format on CPUs and GPUs. Its llama-server program provides an OpenAI-compatible HTTP API (/v1/chat/completions, /v1/completions, /v1/embeddings) and a small built-in web interface. This guide builds llama.cpp from source with CMake, for CPU or with NVIDIA CUDA, tests llama-server with a model downloaded from Hugging Face, runs it as a systemd service under a dedicated user on 127.0.0.1 with an API key, and publishes it over HTTPS with Caddy. A Docker alternative with the official container images is included.

Prerequisites

  • A server running Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12 or Debian 13.
  • A non-root user with sudo rights; see Secure a new Linux server and Set up SSH keys.
  • Optional, for GPU offload: an NVIDIA GPU with a working driver (nvidia-smi prints your GPU). Step 1 of How to install vLLM shows how to install the driver from NVIDIA's repository; see GPU servers for suitable hardware.
  • A domain such as llm.example.com pointing at the server and Caddy installed from Caddy as a reverse proxy, if clients outside the server need the API.
ResourceMinimum (official)Suggested starting point
CPUNot published4 vCPU; generation uses the physical cores
RAMNot publishedAt least the GGUF file size plus 2 to 4 GB for context and the system
GPU (optional)NVIDIA GPU with the CUDA toolkit installed for the CUDA buildEnough VRAM to hold the model so all layers run on the GPU
DiskNot published20 GB for the source, build and a few small models

The llama.cpp project does not publish minimum requirements. The suggested values are a conservative starting point based on the rule of thumb that the model file has to fit in memory, not a benchmark. Larger context sizes and more parallel slots need more memory.

Step 1 — Install the build tools

llama.cpp needs a C/C++ compiler, CMake and Git. Install the OpenSSL development package too: llama.cpp's HTTPS support, which -hf downloads from Hugging Face use, is built with OpenSSL and is enabled by default when the library is present.

Bash
sudo apt update
sudo apt install build-essential cmake git libssl-dev

Clone the repository into /opt/llama.cpp, owned by your admin user:

Bash
sudo mkdir -p /opt/llama.cpp && sudo chown $USER:$USER /opt/llama.cpp
git clone https://github.com/ggml-org/llama.cpp /opt/llama.cpp
cd /opt/llama.cpp

Step 2 — Build llama.cpp

Pick the CPU build, or the CUDA build if the server has an NVIDIA GPU. The CUDA build needs the CUDA toolkit, which provides the nvcc compiler; it comes from the same NVIDIA repository as the driver.

CPU

Bash
cd /opt/llama.cpp
cmake -B build
cmake --build build --config Release -j $(nproc)

NVIDIA CUDA

Bash
sudo apt install cuda-toolkit
export PATH=/usr/local/cuda/bin:$PATH
nvcc --version
cd /opt/llama.cpp
cmake -B build -DGGML_CUDA=ON
cmake --build build --config Release -j $(nproc)

The build takes several minutes. For CUDA builds, the build documentation lets you limit compilation to your GPU's compute capability, for example -DCMAKE_CUDA_ARCHITECTURES="86;89", which shortens the build. NVIDIA's installation guide places the toolkit in a versioned directory such as /usr/local/cuda-13.4; /usr/local/cuda normally points to the installed version. Binaries end up in build/bin. Check that the server program runs:

Bash
./build/bin/llama-server --version

Step 3 — Run a first test with a GGUF model

llama-server loads a model either from a local GGUF file with -m or straight from Hugging Face with -hf user/model[:quant]; without a quant tag it picks Q4_K_M, or the first file in the repository. The project's README uses the small ggml-org/Qwen3.5-0.8B-GGUF model as an example, which is ideal for a first test:

Bash
cd /opt/llama.cpp
./build/bin/llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF --host 127.0.0.1 --port 8080 -c 8192

The server downloads the model into your Hugging Face cache (~/.cache/huggingface/hub), loads it and prints that it is listening on 127.0.0.1:8080. In a second SSH session, test the endpoints:

Bash
curl http://127.0.0.1:8080/health
curl http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"messages": [{"role": "user", "content": "Write a haiku about servers"}]}'

/health returns {"status":"ok"} once the model is loaded (503 while loading), and the chat request returns a JSON completion. Stop the test server with Ctrl+C. The most important options:

OptionDefaultPurpose
-m, --modelnonePath to a local GGUF file
-hf, --hf-repononeDownload a model from Hugging Face as user/model[:quant]
-c, --ctx-size0 (taken from the model)Context size in tokens; set it explicitly to cap memory use
-t, --threads-1 (automatic)CPU threads used for generation
-ngl, --n-gpu-layersautoLayers stored in VRAM: a number, auto or all
-np, --parallel-1 (automatic)Number of server slots for parallel requests
--host, --port127.0.0.1, 8080Listen address and port
--api-keynoneAPI key or comma-separated keys; also LLAMA_API_KEY
--no-webuiweb UI onDisables the built-in web interface

Every option also has an environment variable (for example LLAMA_ARG_CTX_SIZE), listed in the server README.

Step 4 — Create a service user and the API key

Run the server under its own unprivileged user with its own model directory, and keep the API key in a root-only environment file so it does not show up in the process list:

Bash
sudo useradd --system --create-home --home-dir /var/lib/llama --shell /usr/sbin/nologin llama
sudo install -d -o llama -g llama /var/lib/llama/models
echo "LLAMA_API_KEY=$(openssl rand -hex 32)" | sudo tee /etc/llama-server.env > /dev/null
echo "LLAMA_CACHE=/var/lib/llama/models" | sudo tee -a /etc/llama-server.env > /dev/null
sudo chmod 600 /etc/llama-server.env
sudo cat /etc/llama-server.env

LLAMA_CACHE tells llama.cpp where to store -hf downloads, so the service keeps its models in /var/lib/llama/models. Note the API key; clients need it. The service downloads the model again into its own directory, so you can remove the test copy from ~/.cache/huggingface/hub afterwards.

Step 5 — Run llama-server with systemd

Create /etc/systemd/system/llama-server.service:

INI
[Unit]
Description=llama.cpp server
After=network-online.target
Wants=network-online.target

[Service]
User=llama
Group=llama
EnvironmentFile=/etc/llama-server.env
ExecStart=/opt/llama.cpp/build/bin/llama-server -hf ggml-org/Qwen3.5-0.8B-GGUF --host 127.0.0.1 --port 8080 -c 8192
Restart=on-failure
RestartSec=5

[Install]
WantedBy=multi-user.target

To serve a GGUF file you copied to the server yourself, put it in /var/lib/llama/models, make the llama user its owner and replace -hf ggml-org/Qwen3.5-0.8B-GGUF with -m /var/lib/llama/models/your-model.gguf. Enable and start the service, and follow the log until the model is loaded:

Bash
sudo systemctl daemon-reload
sudo systemctl enable --now llama-server
sudo journalctl -u llama-server -f

Now check that the API key is enforced. The first request must fail with 401, the second must list the model:

Bash
KEY=$(sudo grep LLAMA_API_KEY /etc/llama-server.env | cut -d= -f2)
curl -i http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/models -H "Authorization: Bearer $KEY"

With a key configured, the API endpoints require it; /health stays public so monitoring can check the server without credentials.

Step 6 — Publish the API over HTTPS with Caddy

Add a site block to /etc/caddy/Caddyfile, reload Caddy and allow only SSH and web traffic in the firewall:

Caddyfile
llm.example.com {
    reverse_proxy 127.0.0.1:8080
}
Bash
sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable

OpenAI-compatible clients now use https://llm.example.com/v1 as the base URL and your key as the API key. Caddy streams the server-sent events of streaming responses without extra settings. If you prefer Nginx, follow Nginx with Certbot and turn off proxy buffering for the API location.

Run llama-server with Docker instead

The project publishes container images on GitHub Container Registry: ghcr.io/ggml-org/llama.cpp:server for CPU, :server-cuda (CUDA 12) and :server-cuda13 for NVIDIA GPUs, plus ROCm, Vulkan and Intel variants. The server image already listens on all container addresses, so publish the port on loopback only. Create /opt/llama-server/compose.yaml and an .env file with LLAMA_API_KEY= and a key from openssl rand -hex 32 (chmod 600 .env):

YAML
services:
  llama-server:
    image: ghcr.io/ggml-org/llama.cpp:server
    command: ["-hf", "ggml-org/Qwen3.5-0.8B-GGUF", "--port", "8080", "-c", "8192"]
    environment:
      LLAMA_API_KEY: "${LLAMA_API_KEY}"
      LLAMA_CACHE: "/models"
    volumes:
      - ./models:/models
    ports:
      - "127.0.0.1:8080:8080"
    restart: unless-stopped
Bash
cd /opt/llama-server
docker compose up -d
docker compose logs -f

For NVIDIA GPUs, use ghcr.io/ggml-org/llama.cpp:server-cuda, install the NVIDIA Container Toolkit (Step 2 of How to install vLLM) and add a deploy.resources.reservations.devices block with driver: nvidia, count: all and capabilities: [gpu] to the service. The documentation notes that the GPU images are built by CI and not tested beyond that. Use either Docker or the systemd service on port 8080, not both.

Back up and restore

The build itself can be recreated at any time. What you need to keep is the configuration and, if downloads are slow or the files are your own, the models:

Bash
sudo mkdir -p /opt/backups
sudo tar czf /opt/backups/llama-server-$(date +%F).tar.gz /etc/llama-server.env /etc/systemd/system/llama-server.service /var/lib/llama/models

The archive contains your API key, so store it as carefully as the key itself. To restore on a new server, build llama.cpp as in Steps 1 and 2, create the llama user from Step 4, then extract the archive and restart the service:

Bash
sudo tar xzf /opt/backups/llama-server-YYYY-MM-DD.tar.gz -C /
sudo chown -R llama:llama /var/lib/llama
sudo systemctl daemon-reload
sudo systemctl enable --now llama-server

Copy the archive off the server as well.

Update llama.cpp

llama.cpp changes quickly and publishes frequent numbered builds on its GitHub releases page; read the release notes for changes to server options before you update. Pull the latest code, rebuild with the same options (CMake keeps them in the build directory) and restart:

Bash
cd /opt/llama.cpp
git pull
cmake -B build
cmake --build build --config Release -j $(nproc)
sudo systemctl restart llama-server

If a new version misbehaves, check out the previous release tag with git checkout and rebuild. For the Docker variant, run docker compose pull and docker compose up -d.

Troubleshooting

CMake cannot find CUDA or nvcc is not found

The CUDA toolkit is missing or not in your PATH. Install cuda-toolkit from NVIDIA's repository, run export PATH=/usr/local/cuda/bin:$PATH and check nvcc --version. Then delete the build directory and configure again with -DGGML_CUDA=ON.

401 Unauthorized from the API

The server has an API key and the request did not send it, or sent a different one. Send Authorization: Bearer followed by the key from /etc/llama-server.env. In OpenAI client libraries, set the key as api_key.

Downloads with -hf fail

llama.cpp was built without HTTPS support, usually because libssl-dev was not installed. Install it, delete the build directory and build again. For gated or private repositories, add HF_TOKEN= with a Hugging Face access token to /etc/llama-server.env.

Out of memory or failed to allocate when loading a model

The model and its context do not fit in RAM or VRAM. Lower -c, choose a smaller quantization (for example :Q4_K_M instead of :Q8_0), use a smaller model, reduce -np, or set -ngl to a number below the model's layer count so part of the model stays in RAM.

The service fails with permission denied

The llama user cannot read the binary or write to its cache. Check with sudo -u llama ls /opt/llama.cpp/build/bin and sudo -u llama touch /var/lib/llama/models/test, and fix ownership with sudo chown -R llama:llama /var/lib/llama.

Next steps

Frequently asked questions

What is the difference between llama.cpp and Ollama?

llama.cpp is the inference engine and llama-server is its HTTP server. You choose the GGUF file, context size, threads and GPU layers yourself. Ollama adds a model library and its own management layer on top. Use llama-server when you want direct control over these settings.

Do I need a GPU for llama-server?

No. llama.cpp runs GGUF models on the CPU. A CUDA build offloads layers to an NVIDIA GPU, and with the default -ngl auto it places as many layers in VRAM as fit.

Is the llama-server API protected?

Only if you set an API key with --api-key, --api-key-file or the LLAMA_API_KEY variable. With a key set, API requests need an Authorization: Bearer header, while the health endpoint stays public for monitoring. Keep the server on 127.0.0.1 behind a reverse proxy in any case.

Where are models downloaded with -hf stored?

In the Hugging Face cache of the user running the server, by default ~/.cache/huggingface/hub. Set LLAMA_CACHE to use a different directory, as this guide does for the service user.

How large a model can my server run?

The llama.cpp project does not publish requirements. As a rule of thumb, you need at least the size of the GGUF file in RAM or VRAM plus memory for the context, so start with a small quantized model and a fixed context size.

Sources

Generar contrasenya

Please confirm