Skip to content

TutorialsAI & LLM

How to install LocalAI with Docker Compose as an OpenAI-compatible API

Run LocalAI with Docker Compose on CPU or NVIDIA GPU, protect it with an API key, install models from the gallery, add HTTPS with Caddy and keep it backed up.

  • Intermediate
  • 35 min read
  • Updated

Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13

This guide is not available in your language yet, so it is shown in English.

On this page
  1. Prerequisites
  2. Step 1 — Choose the image
  3. Step 2 — Create the project directory and API key
  4. Step 3 — Write the Compose file
  5. Step 4 — Start LocalAI and check the API key
  6. Step 5 — Install a model from the gallery
  7. Step 6 — Test the OpenAI-compatible API
  8. Step 7 — Publish LocalAI over HTTPS with Caddy
  9. Optional: user accounts instead of a shared key
  10. Back up and restore
  11. Update LocalAI
  12. Troubleshooting
  13. 401 Unauthorized
  14. Auto-detected mode as legacy
  15. Models run on the CPU although the server has a GPU
  16. A model installation fails or never finishes
  17. docker compose ps shows health: starting for a long time
  18. Next steps

LocalAI is an open-source engine that serves language, image, speech and embedding models through a REST API compatible with OpenAI's. Applications written for the OpenAI API can use it by changing their base URL. It includes a web interface and a model gallery, and it downloads the matching inference backend when you install a model. This guide runs LocalAI with Docker Compose from the official images, requires an API key for every request, installs a model from the gallery, tests the API, publishes it over HTTPS with Caddy and covers backups, updates and troubleshooting.

Prerequisites

ResourceMinimum (official)Suggested starting point
CPUNot published4 vCPU; set threads to the number of physical cores
RAMNot published8 GB for small quantized language models on CPU
GPU (optional)NVIDIA with driver and Container Toolkit, AMD ROCm, Intel or Vulkan images availableEnough VRAM to hold the model you install
DiskNot published50 GB free for images, backends and models

The LocalAI installation documentation does not list system requirements; the suggested values are a conservative starting point, not a benchmark. Each gallery model downloads its own weights, so plan disk space per model.

Step 1 — Choose the image

LocalAI publishes the same tags on Docker Hub (localai/localai) and Quay (quay.io/go-skynet/local-ai). The documentation's examples use Docker Hub:

HardwareImage
CPU onlylocalai/localai:latest
NVIDIA GPU, CUDA 12localai/localai:latest-gpu-nvidia-cuda-12
NVIDIA GPU, CUDA 13localai/localai:latest-gpu-nvidia-cuda-13
AMD GPU (ROCm)localai/localai:latest-gpu-hipblas
Intel GPUlocalai/localai:latest-gpu-intel
Vulkanlocalai/localai:latest-gpu-vulkan

To pin a release, replace latest with a version tag from the LocalAI releases page on GitHub. Older tutorials mention all-in-one (-aio-) images with preconfigured models; the current container documentation lists only the standard images above, so this guide uses them and installs models from the gallery.

Step 2 — Create the project directory and API key

Create the directory with the four folders LocalAI uses (the container paths must be exactly /models, /backends, /configuration and /data) and a key in .env:

Bash
sudo mkdir -p /opt/localai && sudo chown $USER:$USER /opt/localai
cd /opt/localai
mkdir -p models backends configuration data
echo "LOCALAI_API_KEY=$(openssl rand -hex 32)" > .env
chmod 600 .env

LOCALAI_API_KEY accepts one key or a comma-separated list. Legacy API keys grant full administrative access, so treat the key like a root password.

Step 3 — Write the Compose file

Create /opt/localai/compose.yaml. Files outside these four mounted folders are lost when the container is recreated, so everything LocalAI should keep goes into them:

YAML
services:
  localai:
    image: localai/localai:latest
    environment:
      LOCALAI_API_KEY: "${LOCALAI_API_KEY}"
    volumes:
      - ./models:/models
      - ./backends:/backends
      - ./configuration:/configuration
      - ./data:/data
    ports:
      - "127.0.0.1:8080:8080"
    restart: unless-stopped

For an NVIDIA GPU, change the image to localai/localai:latest-gpu-nvidia-cuda-12 (or the CUDA 13 tag) and add a GPU reservation to the service, at the same indentation as image::

YAML
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]

Step 4 — Start LocalAI and check the API key

Bash
cd /opt/localai
docker compose up -d
docker compose logs -f localai

The log prints CPU feature detection and then the API address. Press Ctrl+C to stop following it. Load the key into your shell and check that the server answers and enforces the key:

Bash
export LOCALAI_API_KEY=$(grep LOCALAI_API_KEY .env | cut -d= -f2)
curl -i http://127.0.0.1:8080/readyz
curl -i http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/models -H "Authorization: Bearer $LOCALAI_API_KEY"

/readyz is a public health check and returns 200. The second request returns 401 with a WWW-Authenticate: Bearer header, and the third returns an empty model list. Clients can send the key as a bearer token or in an x-api-key header.

The easiest way is the web interface. Open an SSH tunnel from your computer, browse to http://localhost:8080 and enter your API key on the sign-in screen:

Bash
ssh -L 8080:127.0.0.1:8080 user@203.0.113.10

Go to Models, then Explore, search for a model such as qwen3-4b (the model used in LocalAI's quick start) and click Install. You can also browse the gallery at models.localai.io. To install from the command line instead, call the gallery API with the gallery name and model name:

Bash
curl http://127.0.0.1:8080/models/apply \
  -H "Authorization: Bearer $LOCALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"id": "localai@qwen3-4b"}'

The response contains a job uuid. Check the job until it reports "processed":true, replacing JOB_ID with that value:

Bash
curl http://127.0.0.1:8080/models/jobs/JOB_ID -H "Authorization: Bearer $LOCALAI_API_KEY"

The first installation also downloads the inference backend the model needs into /opt/localai/backends, so it takes longer than later ones.

Step 6 — Test the OpenAI-compatible API

Send a chat request to the installed model:

Bash
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $LOCALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-4b", "messages": [{"role": "user", "content": "Hello!"}]}'

The response contains the answer in choices[0].message.content. The first request loads the model into memory and is slower than the following ones. OpenAI client libraries work with http://127.0.0.1:8080/v1 as the base URL and your key as the API key.

Step 7 — Publish LocalAI over HTTPS with Caddy

Add a site block to /etc/caddy/Caddyfile, then reload Caddy and allow only SSH and web traffic:

Caddyfile
ai.example.com {
    reverse_proxy 127.0.0.1:8080
}
Bash
sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable

Clients now use https://ai.example.com/v1 as the base URL. Caddy proxies the web interface, streaming responses and WebSocket connections without extra settings.

Optional: user accounts instead of a shared key

For teams, LocalAI offers multi-user mode with admin and user roles, per-user API keys and usage tracking. Add LOCALAI_AUTH: "true" to the environment and run docker compose up -d. The first user to sign in is automatically made administrator, so create that account through the SSH tunnel before you publish the site, or set LOCALAI_ADMIN_EMAIL to the address that should be promoted. New registrations wait for approval by default (LOCALAI_REGISTRATION_MODE=approval). The user database is stored in /data/database.db, which is inside the mounted data folder.

Back up and restore

Everything LocalAI keeps lives in /opt/localai: installed model files and their configuration in models, API key and settings files in configuration, the user database and job state in data, plus compose.yaml and .env. Backends can be downloaded again. Stop the container so the database is consistent, then archive the directory:

Bash
sudo mkdir -p /opt/backups
cd /opt/localai
docker compose stop
sudo tar czf /opt/backups/localai-$(date +%F).tar.gz --exclude=./backends -C /opt/localai .
docker compose start

Model files can be large. If you would rather reinstall models from the gallery after a restore, also add --exclude=./models. To restore, extract the archive into an empty /opt/localai and start the stack:

Bash
sudo mkdir -p /opt/localai
sudo tar xzf /opt/backups/localai-YYYY-MM-DD.tar.gz -C /opt/localai
cd /opt/localai
docker compose up -d

Copy the archives off the server; they contain your API key.

Update LocalAI

Read the release notes on the LocalAI releases page first; version 4.9, for example, changed authentication to deny by default. Back up, then pull the new image and recreate the container. Data in the mounted folders is kept:

Bash
cd /opt/localai
docker compose pull
docker compose up -d
docker image prune

If you pinned a version tag, change it in compose.yaml before pulling.

Troubleshooting

401 Unauthorized

The request did not carry a valid key. Since version 4.9 every route that is not explicitly public requires the key, including /v1/models, /version and the URLs of generated images and audio. Send Authorization: Bearer with the key, an x-api-key header or the key as API key in your client library.

Auto-detected mode as legacy

This NVIDIA error comes from the container runtime configuration, not from LocalAI. Check that sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi works, repeat sudo nvidia-ctk runtime configure --runtime=docker and restart Docker. The LocalAI documentation suggests switching to CDI (driver: nvidia.com/gpu) as the first fix.

Models run on the CPU although the server has a GPU

The stack uses the CPU image or the GPU reservation is missing. Check the image: line, the deploy block and the output of docker compose exec localai nvidia-smi. Recreate the container with docker compose up -d after changes.

A model installation fails or never finishes

Read docker compose logs localai. Typical causes are a full disk (df -h /opt/localai), blocked outbound HTTPS to the gallery or Hugging Face, or a typo in the model ID. Remove a failed model in the web interface and install it again.

docker compose ps shows health: starting for a long time

The image's health check allows a long start period while models and backends download. As long as /readyz answers, the server works; the status changes to healthy once the check passes.

Next steps

Frequently asked questions

What is the difference between LocalAI and Ollama?

Both run open models locally. LocalAI focuses on an OpenAI-compatible API for several model types, including text, images and audio, and loads inference backends on demand from its gallery. Ollama focuses on language models with its own library and CLI.

Does LocalAI require an API key?

Only if you configure one. Without LOCALAI_API_KEY or user authentication, the server does not restrict requests. Since version 4.9, once a key is set, every route that is not explicitly public requires it, including /v1/models and /version.

Which image should I use, and what happened to the AIO images?

The current container documentation lists standard images: localai/localai:latest for CPU and tags such as latest-gpu-nvidia-cuda-12 or latest-gpu-nvidia-cuda-13 for NVIDIA GPUs. They start without models; you install the models you need from the gallery.

Where does LocalAI store models and data?

In the container paths /models, /backends, /configuration and /data. This guide mounts them from /opt/localai so they survive container updates and can be backed up with tar.

Can several people use LocalAI with their own accounts?

Yes. Set LOCALAI_AUTH=true for multi-user mode with admin and user roles and per-user API keys. The first user to sign in becomes the administrator, so complete that step before you publish the site.

Gerar senha

Please confirm