How to install LocalAI with Docker Compose as an OpenAI-compatible API
Run LocalAI with Docker Compose on CPU or NVIDIA GPU, protect it with an API key, install models from the gallery, add HTTPS with Caddy and keep it backed up.
- Intermediate
- 35 min read
- Updated
Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13
This guide is not available in your language yet, so it is shown in English.
On this page
- Prerequisites
- Step 1 — Choose the image
- Step 2 — Create the project directory and API key
- Step 3 — Write the Compose file
- Step 4 — Start LocalAI and check the API key
- Step 5 — Install a model from the gallery
- Step 6 — Test the OpenAI-compatible API
- Step 7 — Publish LocalAI over HTTPS with Caddy
- Optional: user accounts instead of a shared key
- Back up and restore
- Update LocalAI
- Troubleshooting
- 401 Unauthorized
- Auto-detected mode as legacy
- Models run on the CPU although the server has a GPU
- A model installation fails or never finishes
- docker compose ps shows health: starting for a long time
- Next steps
LocalAI is an open-source engine that serves language, image, speech and embedding models through a REST API compatible with OpenAI's. Applications written for the OpenAI API can use it by changing their base URL. It includes a web interface and a model gallery, and it downloads the matching inference backend when you install a model. This guide runs LocalAI with Docker Compose from the official images, requires an API key for every request, installs a model from the gallery, tests the API, publishes it over HTTPS with Caddy and covers backups, updates and troubleshooting.
Prerequisites
- A server running Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12 or Debian 13.
- A non-root user with
sudorights; see Secure a new Linux server and Set up SSH keys. - Docker Engine with the Compose plugin from Install Docker on Ubuntu or Install Docker on Debian. The LocalAI documentation recommends the container method.
- Optional, for NVIDIA GPUs: the NVIDIA driver and the NVIDIA Container Toolkit, installed as in Steps 1 and 2 of How to install vLLM. See GPU servers for suitable hardware.
- A domain such as
ai.example.compointing at the server, and Caddy from Caddy as a reverse proxy.
| Resource | Minimum (official) | Suggested starting point |
|---|---|---|
| CPU | Not published | 4 vCPU; set threads to the number of physical cores |
| RAM | Not published | 8 GB for small quantized language models on CPU |
| GPU (optional) | NVIDIA with driver and Container Toolkit, AMD ROCm, Intel or Vulkan images available | Enough VRAM to hold the model you install |
| Disk | Not published | 50 GB free for images, backends and models |
The LocalAI installation documentation does not list system requirements; the suggested values are a conservative starting point, not a benchmark. Each gallery model downloads its own weights, so plan disk space per model.
Step 1 — Choose the image
LocalAI publishes the same tags on Docker Hub (localai/localai) and Quay (quay.io/go-skynet/local-ai). The documentation's examples use Docker Hub:
| Hardware | Image |
|---|---|
| CPU only | localai/localai:latest |
| NVIDIA GPU, CUDA 12 | localai/localai:latest-gpu-nvidia-cuda-12 |
| NVIDIA GPU, CUDA 13 | localai/localai:latest-gpu-nvidia-cuda-13 |
| AMD GPU (ROCm) | localai/localai:latest-gpu-hipblas |
| Intel GPU | localai/localai:latest-gpu-intel |
| Vulkan | localai/localai:latest-gpu-vulkan |
To pin a release, replace latest with a version tag from the LocalAI releases page on GitHub. Older tutorials mention all-in-one (-aio-) images with preconfigured models; the current container documentation lists only the standard images above, so this guide uses them and installs models from the gallery.
Step 2 — Create the project directory and API key
Create the directory with the four folders LocalAI uses (the container paths must be exactly /models, /backends, /configuration and /data) and a key in .env:
sudo mkdir -p /opt/localai && sudo chown $USER:$USER /opt/localai
cd /opt/localai
mkdir -p models backends configuration data
echo "LOCALAI_API_KEY=$(openssl rand -hex 32)" > .env
chmod 600 .envLOCALAI_API_KEY accepts one key or a comma-separated list. Legacy API keys grant full administrative access, so treat the key like a root password.
Step 3 — Write the Compose file
Create /opt/localai/compose.yaml. Files outside these four mounted folders are lost when the container is recreated, so everything LocalAI should keep goes into them:
services:
localai:
image: localai/localai:latest
environment:
LOCALAI_API_KEY: "${LOCALAI_API_KEY}"
volumes:
- ./models:/models
- ./backends:/backends
- ./configuration:/configuration
- ./data:/data
ports:
- "127.0.0.1:8080:8080"
restart: unless-stoppedFor an NVIDIA GPU, change the image to localai/localai:latest-gpu-nvidia-cuda-12 (or the CUDA 13 tag) and add a GPU reservation to the service, at the same indentation as image::
deploy:
resources:
reservations:
devices:
- driver: nvidia
count: all
capabilities: [gpu]Step 4 — Start LocalAI and check the API key
cd /opt/localai
docker compose up -d
docker compose logs -f localaiThe log prints CPU feature detection and then the API address. Press Ctrl+C to stop following it. Load the key into your shell and check that the server answers and enforces the key:
export LOCALAI_API_KEY=$(grep LOCALAI_API_KEY .env | cut -d= -f2)
curl -i http://127.0.0.1:8080/readyz
curl -i http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/models -H "Authorization: Bearer $LOCALAI_API_KEY"/readyz is a public health check and returns 200. The second request returns 401 with a WWW-Authenticate: Bearer header, and the third returns an empty model list. Clients can send the key as a bearer token or in an x-api-key header.
Step 5 — Install a model from the gallery
The easiest way is the web interface. Open an SSH tunnel from your computer, browse to http://localhost:8080 and enter your API key on the sign-in screen:
ssh -L 8080:127.0.0.1:8080 user@203.0.113.10Go to Models, then Explore, search for a model such as qwen3-4b (the model used in LocalAI's quick start) and click Install. You can also browse the gallery at models.localai.io. To install from the command line instead, call the gallery API with the gallery name and model name:
curl http://127.0.0.1:8080/models/apply \
-H "Authorization: Bearer $LOCALAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"id": "localai@qwen3-4b"}'The response contains a job uuid. Check the job until it reports "processed":true, replacing JOB_ID with that value:
curl http://127.0.0.1:8080/models/jobs/JOB_ID -H "Authorization: Bearer $LOCALAI_API_KEY"The first installation also downloads the inference backend the model needs into /opt/localai/backends, so it takes longer than later ones.
Step 6 — Test the OpenAI-compatible API
Send a chat request to the installed model:
curl http://127.0.0.1:8080/v1/chat/completions \
-H "Authorization: Bearer $LOCALAI_API_KEY" \
-H "Content-Type: application/json" \
-d '{"model": "qwen3-4b", "messages": [{"role": "user", "content": "Hello!"}]}'The response contains the answer in choices[0].message.content. The first request loads the model into memory and is slower than the following ones. OpenAI client libraries work with http://127.0.0.1:8080/v1 as the base URL and your key as the API key.
Step 7 — Publish LocalAI over HTTPS with Caddy
Add a site block to /etc/caddy/Caddyfile, then reload Caddy and allow only SSH and web traffic:
ai.example.com {
reverse_proxy 127.0.0.1:8080
}sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enableClients now use https://ai.example.com/v1 as the base URL. Caddy proxies the web interface, streaming responses and WebSocket connections without extra settings.
Optional: user accounts instead of a shared key
For teams, LocalAI offers multi-user mode with admin and user roles, per-user API keys and usage tracking. Add LOCALAI_AUTH: "true" to the environment and run docker compose up -d. The first user to sign in is automatically made administrator, so create that account through the SSH tunnel before you publish the site, or set LOCALAI_ADMIN_EMAIL to the address that should be promoted. New registrations wait for approval by default (LOCALAI_REGISTRATION_MODE=approval). The user database is stored in /data/database.db, which is inside the mounted data folder.
Back up and restore
Everything LocalAI keeps lives in /opt/localai: installed model files and their configuration in models, API key and settings files in configuration, the user database and job state in data, plus compose.yaml and .env. Backends can be downloaded again. Stop the container so the database is consistent, then archive the directory:
sudo mkdir -p /opt/backups
cd /opt/localai
docker compose stop
sudo tar czf /opt/backups/localai-$(date +%F).tar.gz --exclude=./backends -C /opt/localai .
docker compose startModel files can be large. If you would rather reinstall models from the gallery after a restore, also add --exclude=./models. To restore, extract the archive into an empty /opt/localai and start the stack:
sudo mkdir -p /opt/localai
sudo tar xzf /opt/backups/localai-YYYY-MM-DD.tar.gz -C /opt/localai
cd /opt/localai
docker compose up -dCopy the archives off the server; they contain your API key.
Update LocalAI
Read the release notes on the LocalAI releases page first; version 4.9, for example, changed authentication to deny by default. Back up, then pull the new image and recreate the container. Data in the mounted folders is kept:
cd /opt/localai
docker compose pull
docker compose up -d
docker image pruneIf you pinned a version tag, change it in compose.yaml before pulling.
Troubleshooting
401 Unauthorized
The request did not carry a valid key. Since version 4.9 every route that is not explicitly public requires the key, including /v1/models, /version and the URLs of generated images and audio. Send Authorization: Bearer with the key, an x-api-key header or the key as API key in your client library.
Auto-detected mode as legacy
This NVIDIA error comes from the container runtime configuration, not from LocalAI. Check that sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi works, repeat sudo nvidia-ctk runtime configure --runtime=docker and restart Docker. The LocalAI documentation suggests switching to CDI (driver: nvidia.com/gpu) as the first fix.
Models run on the CPU although the server has a GPU
The stack uses the CPU image or the GPU reservation is missing. Check the image: line, the deploy block and the output of docker compose exec localai nvidia-smi. Recreate the container with docker compose up -d after changes.
A model installation fails or never finishes
Read docker compose logs localai. Typical causes are a full disk (df -h /opt/localai), blocked outbound HTTPS to the gallery or Hugging Face, or a typo in the model ID. Remove a failed model in the web interface and install it again.
docker compose ps shows health: starting for a long time
The image's health check allows a long start period while models and backends download. As long as /readyz answers, the server works; the status changes to healthy once the check passes.
Next steps
- Use LocalAI as an OpenAI-compatible backend for Open WebUI.
- Compare with Ollama, llama.cpp server and vLLM.
- Choose a server on the LLM API hosting page.
- Read the official documentation at https://localai.io for backends, galleries and every setting.
Frequently asked questions
What is the difference between LocalAI and Ollama?
Both run open models locally. LocalAI focuses on an OpenAI-compatible API for several model types, including text, images and audio, and loads inference backends on demand from its gallery. Ollama focuses on language models with its own library and CLI.
Does LocalAI require an API key?
Only if you configure one. Without LOCALAI_API_KEY or user authentication, the server does not restrict requests. Since version 4.9, once a key is set, every route that is not explicitly public requires it, including /v1/models and /version.
Which image should I use, and what happened to the AIO images?
The current container documentation lists standard images: localai/localai:latest for CPU and tags such as latest-gpu-nvidia-cuda-12 or latest-gpu-nvidia-cuda-13 for NVIDIA GPUs. They start without models; you install the models you need from the gallery.
Where does LocalAI store models and data?
In the container paths /models, /backends, /configuration and /data. This guide mounts them from /opt/localai so they survive container updates and can be backed up with tar.
Can several people use LocalAI with their own accounts?
Yes. Set LOCALAI_AUTH=true for multi-user mode with admin and user roles and per-user API keys. The first user to sign in becomes the administrator, so complete that step before you publish the site.
Sources
- localai.io/docs/installation
- localai.io/docs/installation/containers
- localai.io/docs/basics/getting_started
- localai.io/docs/getting-started/models
- localai.io/docs/models
- localai.io/docs/features/authentication
- localai.io/docs/reference/cli-reference
- localai.io/docs/advanced
- raw.githubusercontent.com/mudler/LocalAI/master/README.md
- raw.githubusercontent.com/mudler/LocalAI/master/Dockerfile
- raw.githubusercontent.com/mudler/LocalAI/master/entrypoint.sh
- docs.docker.com/compose/how-tos/gpu-support