# How to install LocalAI with Docker Compose as an OpenAI-compatible API

> Run LocalAI with Docker Compose on CPU or NVIDIA GPU, protect it with an API key, install models from the gallery, add HTTPS with Caddy and keep it backed up.

Difficulty: Intermediate\
Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13

LocalAI is an open-source engine that serves language, image, speech and embedding models through a REST API compatible with OpenAI's. Applications written for the OpenAI API can use it by changing their base URL. It includes a web interface and a model gallery, and it downloads the matching inference backend when you install a model. This guide runs **LocalAI with Docker Compose** from the official images, requires an **API key** for every request, installs a model from the gallery, tests the API, publishes it over **HTTPS with Caddy** and covers backups, updates and troubleshooting.

## Prerequisites

- A server running **Ubuntu 24.04 LTS**, **Ubuntu 26.04 LTS**, **Debian 12** or **Debian 13**.
- A non-root user with `sudo` rights; see [Secure a new Linux server](/guides/secure-a-new-linux-server) and [Set up SSH keys](/guides/ssh-keys).
- Docker Engine with the Compose plugin from [Install Docker on Ubuntu](/guides/install-docker-ubuntu) or [Install Docker on Debian](/guides/install-docker-debian). The LocalAI documentation recommends the container method.
- Optional, for NVIDIA GPUs: the NVIDIA driver and the NVIDIA Container Toolkit, installed as in Steps 1 and 2 of [How to install vLLM](/guides/install-vllm). See [GPU servers](/gpu-servers) for suitable hardware.
- A domain such as `ai.example.com` pointing at the server, and Caddy from [Caddy as a reverse proxy](/guides/caddy-reverse-proxy).

| Resource | Minimum (official) | Suggested starting point |
|---|---|---|
| CPU | Not published | 4 vCPU; set threads to the number of physical cores |
| RAM | Not published | 8 GB for small quantized language models on CPU |
| GPU (optional) | NVIDIA with driver and Container Toolkit, AMD ROCm, Intel or Vulkan images available | Enough VRAM to hold the model you install |
| Disk | Not published | 50 GB free for images, backends and models |

The LocalAI installation documentation does not list system requirements; the suggested values are a conservative starting point, not a benchmark. Each gallery model downloads its own weights, so plan disk space per model.

## Step 1 — Choose the image

LocalAI publishes the same tags on Docker Hub (`localai/localai`) and Quay (`quay.io/go-skynet/local-ai`). The documentation's examples use Docker Hub:

| Hardware | Image |
|---|---|
| CPU only | `localai/localai:latest` |
| NVIDIA GPU, CUDA 12 | `localai/localai:latest-gpu-nvidia-cuda-12` |
| NVIDIA GPU, CUDA 13 | `localai/localai:latest-gpu-nvidia-cuda-13` |
| AMD GPU (ROCm) | `localai/localai:latest-gpu-hipblas` |
| Intel GPU | `localai/localai:latest-gpu-intel` |
| Vulkan | `localai/localai:latest-gpu-vulkan` |

To pin a release, replace `latest` with a version tag from the LocalAI releases page on GitHub. Older tutorials mention all-in-one (`-aio-`) images with preconfigured models; the current container documentation lists only the standard images above, so this guide uses them and installs models from the gallery.

## Step 2 — Create the project directory and API key

Create the directory with the four folders LocalAI uses (the container paths must be exactly `/models`, `/backends`, `/configuration` and `/data`) and a key in `.env`:

```bash
sudo mkdir -p /opt/localai && sudo chown $USER:$USER /opt/localai
cd /opt/localai
mkdir -p models backends configuration data
echo "LOCALAI_API_KEY=$(openssl rand -hex 32)" > .env
chmod 600 .env
```

`LOCALAI_API_KEY` accepts one key or a comma-separated list. Legacy API keys grant full administrative access, so treat the key like a root password.

## Step 3 — Write the Compose file

Create `/opt/localai/compose.yaml`. Files outside these four mounted folders are lost when the container is recreated, so everything LocalAI should keep goes into them:

```yaml
services:
  localai:
    image: localai/localai:latest
    environment:
      LOCALAI_API_KEY: "${LOCALAI_API_KEY}"
    volumes:
      - ./models:/models
      - ./backends:/backends
      - ./configuration:/configuration
      - ./data:/data
    ports:
      - "127.0.0.1:8080:8080"
    restart: unless-stopped
```

For an NVIDIA GPU, change the image to `localai/localai:latest-gpu-nvidia-cuda-12` (or the CUDA 13 tag) and add a GPU reservation to the service, at the same indentation as `image:`:

```yaml
    deploy:
      resources:
        reservations:
          devices:
            - driver: nvidia
              count: all
              capabilities: [gpu]
```

> **Note**
>
> LocalAI's own Compose example uses `driver: nvidia.com/gpu`, which relies on CDI and is recommended for NVIDIA Container Toolkit 1.14 or newer. `driver: nvidia` works with the runtime configured by `nvidia-ctk runtime configure`. If one does not work on your system, try the other.

## Step 4 — Start LocalAI and check the API key

```bash
cd /opt/localai
docker compose up -d
docker compose logs -f localai
```

The log prints CPU feature detection and then the API address. Press `Ctrl+C` to stop following it. Load the key into your shell and check that the server answers and enforces the key:

```bash
export LOCALAI_API_KEY=$(grep LOCALAI_API_KEY .env | cut -d= -f2)
curl -i http://127.0.0.1:8080/readyz
curl -i http://127.0.0.1:8080/v1/models
curl http://127.0.0.1:8080/v1/models -H "Authorization: Bearer $LOCALAI_API_KEY"
```

`/readyz` is a public health check and returns `200`. The second request returns `401` with a `WWW-Authenticate: Bearer` header, and the third returns an empty model list. Clients can send the key as a bearer token or in an `x-api-key` header.

## Step 5 — Install a model from the gallery

The easiest way is the web interface. Open an SSH tunnel from your computer, browse to `http://localhost:8080` and enter your API key on the sign-in screen:

```bash
ssh -L 8080:127.0.0.1:8080 user@203.0.113.10
```

Go to **Models**, then **Explore**, search for a model such as `qwen3-4b` (the model used in LocalAI's quick start) and click **Install**. You can also browse the gallery at `models.localai.io`. To install from the command line instead, call the gallery API with the gallery name and model name:

```bash
curl http://127.0.0.1:8080/models/apply \
  -H "Authorization: Bearer $LOCALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"id": "localai@qwen3-4b"}'
```

The response contains a job `uuid`. Check the job until it reports `"processed":true`, replacing `JOB_ID` with that value:

```bash
curl http://127.0.0.1:8080/models/jobs/JOB_ID -H "Authorization: Bearer $LOCALAI_API_KEY"
```

The first installation also downloads the inference backend the model needs into `/opt/localai/backends`, so it takes longer than later ones.

## Step 6 — Test the OpenAI-compatible API

Send a chat request to the installed model:

```bash
curl http://127.0.0.1:8080/v1/chat/completions \
  -H "Authorization: Bearer $LOCALAI_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "qwen3-4b", "messages": [{"role": "user", "content": "Hello!"}]}'
```

The response contains the answer in `choices[0].message.content`. The first request loads the model into memory and is slower than the following ones. OpenAI client libraries work with `http://127.0.0.1:8080/v1` as the base URL and your key as the API key.

## Step 7 — Publish LocalAI over HTTPS with Caddy

Add a site block to `/etc/caddy/Caddyfile`, then reload Caddy and allow only SSH and web traffic:

```caddyfile
ai.example.com {
    reverse_proxy 127.0.0.1:8080
}
```

```bash
sudo systemctl reload caddy
sudo ufw allow OpenSSH
sudo ufw allow 80/tcp
sudo ufw allow 443/tcp
sudo ufw enable
```

Clients now use `https://ai.example.com/v1` as the base URL. Caddy proxies the web interface, streaming responses and WebSocket connections without extra settings.

> **Warning**
>
> Never publish port 8080 as `"8080:8080"` and never run LocalAI without `LOCALAI_API_KEY` or user authentication on a reachable address. Docker-published ports bypass ufw, and without a key the API accepts every request, including model installation.

### Optional: user accounts instead of a shared key

For teams, LocalAI offers multi-user mode with admin and user roles, per-user API keys and usage tracking. Add `LOCALAI_AUTH: "true"` to the environment and run `docker compose up -d`. The first user to sign in is automatically made administrator, so create that account through the SSH tunnel before you publish the site, or set `LOCALAI_ADMIN_EMAIL` to the address that should be promoted. New registrations wait for approval by default (`LOCALAI_REGISTRATION_MODE=approval`). The user database is stored in `/data/database.db`, which is inside the mounted `data` folder.

## Back up and restore

Everything LocalAI keeps lives in `/opt/localai`: installed model files and their configuration in `models`, API key and settings files in `configuration`, the user database and job state in `data`, plus `compose.yaml` and `.env`. Backends can be downloaded again. Stop the container so the database is consistent, then archive the directory:

```bash
sudo mkdir -p /opt/backups
cd /opt/localai
docker compose stop
sudo tar czf /opt/backups/localai-$(date +%F).tar.gz --exclude=./backends -C /opt/localai .
docker compose start
```

Model files can be large. If you would rather reinstall models from the gallery after a restore, also add `--exclude=./models`. To restore, extract the archive into an empty `/opt/localai` and start the stack:

```bash
sudo mkdir -p /opt/localai
sudo tar xzf /opt/backups/localai-YYYY-MM-DD.tar.gz -C /opt/localai
cd /opt/localai
docker compose up -d
```

Copy the archives off the server; they contain your API key.

## Update LocalAI

Read the release notes on the LocalAI releases page first; version 4.9, for example, changed authentication to deny by default. Back up, then pull the new image and recreate the container. Data in the mounted folders is kept:

```bash
cd /opt/localai
docker compose pull
docker compose up -d
docker image prune
```

If you pinned a version tag, change it in `compose.yaml` before pulling.

## Troubleshooting

### 401 Unauthorized

The request did not carry a valid key. Since version 4.9 every route that is not explicitly public requires the key, including `/v1/models`, `/version` and the URLs of generated images and audio. Send `Authorization: Bearer` with the key, an `x-api-key` header or the key as API key in your client library.

### Auto-detected mode as legacy

This NVIDIA error comes from the container runtime configuration, not from LocalAI. Check that `sudo docker run --rm --runtime=nvidia --gpus all ubuntu nvidia-smi` works, repeat `sudo nvidia-ctk runtime configure --runtime=docker` and restart Docker. The LocalAI documentation suggests switching to CDI (`driver: nvidia.com/gpu`) as the first fix.

### Models run on the CPU although the server has a GPU

The stack uses the CPU image or the GPU reservation is missing. Check the `image:` line, the `deploy` block and the output of `docker compose exec localai nvidia-smi`. Recreate the container with `docker compose up -d` after changes.

### A model installation fails or never finishes

Read `docker compose logs localai`. Typical causes are a full disk (`df -h /opt/localai`), blocked outbound HTTPS to the gallery or Hugging Face, or a typo in the model ID. Remove a failed model in the web interface and install it again.

### docker compose ps shows health: starting for a long time

The image's health check allows a long start period while models and backends download. As long as `/readyz` answers, the server works; the status changes to healthy once the check passes.

## Next steps

- Use LocalAI as an OpenAI-compatible backend for [Open WebUI](/guides/open-webui-ollama).
- Compare with [Ollama](/guides/install-ollama), [llama.cpp server](/guides/llama-cpp-server) and [vLLM](/guides/install-vllm).
- Choose a server on the [LLM API hosting](/llm-api-hosting) page.
- Read the official documentation at https://localai.io for backends, galleries and every setting.

## Frequently asked questions

### What is the difference between LocalAI and Ollama?

Both run open models locally. LocalAI focuses on an OpenAI-compatible API for several model types, including text, images and audio, and loads inference backends on demand from its gallery. Ollama focuses on language models with its own library and CLI.

### Does LocalAI require an API key?

Only if you configure one. Without LOCALAI_API_KEY or user authentication, the server does not restrict requests. Since version 4.9, once a key is set, every route that is not explicitly public requires it, including /v1/models and /version.

### Which image should I use, and what happened to the AIO images?

The current container documentation lists standard images: localai/localai:latest for CPU and tags such as latest-gpu-nvidia-cuda-12 or latest-gpu-nvidia-cuda-13 for NVIDIA GPUs. They start without models; you install the models you need from the gallery.

### Where does LocalAI store models and data?

In the container paths /models, /backends, /configuration and /data. This guide mounts them from /opt/localai so they survive container updates and can be backed up with tar.

### Can several people use LocalAI with their own accounts?

Yes. Set LOCALAI_AUTH=true for multi-user mode with admin and user roles and per-user API keys. The first user to sign in becomes the administrator, so complete that step before you publish the site.

---

Source: <https://hyperdc.com/guides/tutorials/install-localai>\
Updated: 2026-10-09
