Skip to content

TutorialsAI & LLM

How to install Ollama on Ubuntu or Debian and run local LLMs

Install Ollama on Ubuntu or Debian with the official script or tarball, run models on CPU or an NVIDIA GPU, keep the API private and update it safely.

  • Beginner
  • 30 min read
  • Updated

Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13

This guide is not available in your language yet, so it is shown in English.

On this page
  1. Prerequisites
  2. Step 1 — Prepare the server
  3. Step 2 — Download, review and run the install script
  4. Alternative: manual install from the tarball
  5. Step 3 — Pull and run a model
  6. Step 4 — Test the REST API
  7. Step 5 — Configure the service
  8. Step 6 — Keep the API private and reach it securely
  9. Run Ollama in Docker instead
  10. Back up and restore
  11. Update Ollama
  12. Troubleshooting
  13. Reading the logs
  14. The GPU is not used and ollama ps shows 100% CPU
  15. model requires more system memory than is available
  16. 403 Forbidden through a reverse proxy or tunnel
  17. The ollama command cannot connect to the server
  18. Next steps

Ollama is a runtime that downloads open large language models (LLMs) and serves them through a simple command line and a REST API on port 11434. It is the easiest way to run models such as Llama, Gemma or Qwen on your own server, and many tools, including Open WebUI, speak to it directly. This guide installs Ollama with its official Linux install script, explains exactly what the script changes, and shows the manual tarball method as an alternative. You then run your first model on CPU or GPU, test the API with curl, move model storage, keep the API private and learn how to update, back up and troubleshoot the service.

Prerequisites

  • A server running Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12 or Debian 13 on x86_64 (amd64) or arm64.
  • A non-root user with sudo rights. If you have not set one up yet, follow Secure a new Linux server and Set up SSH keys first.
  • Optional, for GPU acceleration: an NVIDIA GPU with compute capability 5.0 or newer and driver 550 or newer (cards with compute capability 5.0 to 6.2 need driver 570 or newer), or a supported AMD Radeon or Instinct GPU with the ROCm v7 driver. See GPU servers for suitable hardware.
  • Enough disk space for the models you plan to download.
ResourceMinimum (official)Suggested starting point
CPUNot published4 vCPU for small models on CPU only
RAMOllama library guidance: 7B models generally need at least 8 GB, 13B at least 16 GB, 70B at least 64 GB16 GB for models up to about 8B parameters
GPU (optional)NVIDIA compute capability 5.0+ with driver 550+, or AMD with ROCm v7Enough VRAM to hold the whole model, so ollama ps shows 100% GPU
DiskNot published50 GB free; each model page lists its download size

The RAM figures come from Ollama's model library and predate current model families, so treat them as a floor. The suggested values are a conservative starting point, not a benchmark. Memory use grows with the context length and with the number of parallel requests: the documentation states that required memory scales with OLLAMA_NUM_PARALLEL multiplied by OLLAMA_CONTEXT_LENGTH. Without an override, Ollama picks the default context length from the available VRAM: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more.

Step 1 — Prepare the server

The install script needs curl, zstd (the Linux release archives are .tar.zst files) and lspci or lshw to look for a GPU. Install them:

Bash
sudo apt update
sudo apt install curl zstd pciutils

If the server has an NVIDIA GPU, install the driver before Ollama so you control which driver package is used. The commands below follow NVIDIA's driver installation guide for x86_64 servers: they add NVIDIA's network repository and install the nvidia-open driver. On Debian, NVIDIA's guide also requires the contrib component. The guide enables it with add-apt-repository, which Debian 13 no longer ships, so add contrib next to main in your APT sources yourself (the Components: line in /etc/apt/sources.list.d/debian.sources, or the deb lines in /etc/apt/sources.list, depending on which file your system uses) before you run the Debian commands.

Ubuntu

Bash
sudo apt install linux-headers-$(uname -r)
distro=ubuntu$(. /etc/os-release && echo "$VERSION_ID" | tr -d .)
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install nvidia-open
sudo reboot

Debian

Bash
sudo apt install linux-headers-$(uname -r)
distro=debian$(. /etc/os-release && echo "$VERSION_ID")
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt -V install nvidia-open
sudo reboot

After the reboot, nvidia-smi should print a table with your GPU, the driver version and the highest CUDA version the driver supports. If you skip this step, the Ollama script detects an NVIDIA GPU without a working driver and installs CUDA drivers itself on Ubuntu and Debian. For AMD GPUs, install the ROCm v7 driver with AMD's amdgpu-install utility as described in AMD's ROCm documentation.

Step 2 — Download, review and run the install script

Ollama's official Linux installer is a shell script. Download it to a file and read it before you run it:

Bash
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
sh ollama-install.sh

Run the script as your normal sudo user, not with sudo: it calls sudo itself when it needs root, and it adds the user who runs it to the ollama group. The documentation shows the same installer as the one-liner curl -fsSL https://ollama.com/install.sh | sh; downloading first simply lets you review it. The script:

  • downloads the Ollama release archive for your CPU architecture and extracts it under /usr/local (binary in /usr/local/bin, libraries in /usr/local/lib/ollama), removing any older libraries first;
  • creates a system user and group called ollama with the home directory /usr/share/ollama, adds it to the render and video groups when they exist, and adds your user to the ollama group;
  • writes /etc/systemd/system/ollama.service, which runs ollama serve as the ollama user with Restart=always, then enables and starts it;
  • looks for a GPU: it reports an NVIDIA GPU if nvidia-smi works, installs CUDA drivers from NVIDIA's repository if it finds an NVIDIA card without a driver, downloads the extra ROCm libraries for AMD cards, and otherwise warns that Ollama will run in CPU-only mode.

When it finishes, it prints that the Ollama API is available at 127.0.0.1:11434. Verify the service and the binary:

Bash
systemctl status ollama --no-pager
ollama -v
curl http://127.0.0.1:11434/api/tags

The service should be active (running), ollama -v prints the version, and the API returns a JSON list of models, which is empty for now.

Alternative: manual install from the tarball

If you prefer not to run the script, extract the release archive yourself. This variant installs into /usr instead of /usr/local. On arm64 servers use ollama-linux-arm64.tar.zst; for AMD GPUs also extract ollama-linux-amd64-rocm.tar.zst the same way.

Bash
curl -fsSL https://ollama.com/download/ollama-linux-amd64.tar.zst | sudo tar x -C /usr
sudo useradd -r -s /bin/false -U -m -d /usr/share/ollama ollama
sudo usermod -a -G ollama $(whoami)

Then create the service file /etc/systemd/system/ollama.service with this content:

INI
[Unit]
Description=Ollama Service
After=network-online.target

[Service]
ExecStart=/usr/bin/ollama serve
User=ollama
Group=ollama
Restart=always
RestartSec=3
Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"

[Install]
WantedBy=multi-user.target
Bash
sudo systemctl daemon-reload
sudo systemctl enable --now ollama

Step 3 — Pull and run a model

Browse the model library at ollama.com/library and check each model's download size. A small model such as llama3.2 (the 3B variant is a 2.0 GB download) is a good first test even without a GPU:

Bash
ollama pull llama3.2
ollama run llama3.2

ollama run opens an interactive prompt; type a question, and type /bye to leave. You can also pass a prompt directly, for example ollama run llama3.2 "Explain DNS in one sentence". Manage models with these commands:

Bash
ollama list
ollama ps
ollama stop llama3.2
ollama rm llama3.2

ollama list shows downloaded models, and ollama ps shows loaded models. In the ollama ps output, the PROCESSOR column tells you where the model runs: 100% GPU means it fits completely in VRAM, 100% CPU means it runs from system memory, and a split such as 48%/52% CPU/GPU means it did not fit in VRAM. Models unload after 5 minutes without requests by default.

Step 4 — Test the REST API

The API is what front ends and your own code use. Send a generation request without streaming:

Bash
curl http://127.0.0.1:11434/api/generate -d '{
  "model": "llama3.2",
  "prompt": "Why is the sky blue?",
  "stream": false
}'

Ollama also offers an OpenAI-compatible API under /v1, so OpenAI client libraries work by pointing their base URL at http://127.0.0.1:11434/v1. The client needs some API key value, which Ollama ignores:

Bash
curl http://127.0.0.1:11434/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "llama3.2", "messages": [{"role": "user", "content": "Say this is a test"}]}'

You should see a JSON response with the model's answer in response (native API) or choices[0].message.content (OpenAI API).

Step 5 — Configure the service

Ollama reads its settings from environment variables. Set them for the systemd service with an override file rather than editing the unit, so that updates do not overwrite your changes:

Bash
sudo systemctl edit ollama

Add the variables you need in the [Service] section of the editor that opens, for example:

INI
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_KEEP_ALIVE=10m"
Bash
sudo systemctl daemon-reload
sudo systemctl restart ollama

The variables you are most likely to need:

VariableDefaultWhat it does
OLLAMA_HOST127.0.0.1:11434Address and port the API listens on
OLLAMA_MODELS/usr/share/ollama/.ollama/modelsModel storage directory
OLLAMA_CONTEXT_LENGTHBased on VRAM (4k below 24 GiB)Default context window in tokens
OLLAMA_KEEP_ALIVE5mHow long a model stays loaded after the last request
OLLAMA_NUM_PARALLEL1Parallel requests per model
OLLAMA_MAX_LOADED_MODELS3 per GPU, or 3 on CPUModels kept in memory at the same time

To keep models on a larger disk, stop the service, copy the existing models, give the ollama user ownership and point OLLAMA_MODELS at the new directory with sudo systemctl edit ollama:

Bash
sudo systemctl stop ollama
sudo mkdir -p /srv/ollama/models
sudo cp -a /usr/share/ollama/.ollama/models/. /srv/ollama/models/
sudo chown -R ollama:ollama /srv/ollama

Add Environment="OLLAMA_MODELS=/srv/ollama/models" to the override, then start Ollama again and confirm with ollama list that your models are still there.

Step 6 — Keep the API private and reach it securely

Ollama has no authentication for local requests. Anyone who can reach port 11434 can run and delete models, so leave OLLAMA_HOST at its default 127.0.0.1:11434 and choose one of these patterns instead of listening on every address:

  • From your own computer, open an SSH tunnel and use http://localhost:11434 locally: ssh -L 11434:127.0.0.1:11434 [email protected].
  • For a web chat interface, run Open WebUI on the same server. How to run Open WebUI with Ollama shows a setup that reaches Ollama on 127.0.0.1 without opening the port.
  • For remote API clients, put Caddy in front of Ollama and require a token. Generate one with openssl rand -hex 32, then add a site block:
Caddyfile
ollama.example.com {
    @authorized header Authorization "Bearer change-me-long-random-token"
    handle @authorized {
        reverse_proxy 127.0.0.1:11434 {
            header_up Host localhost:11434
        }
    }
    handle {
        respond 401
    }
}

Reload Caddy with sudo systemctl reload caddy. The header_up Host line matters: Ollama's documentation rewrites the Host header to localhost:11434 in its proxy examples, and requests with a foreign host name are rejected otherwise. Clients then send Authorization: Bearer with your token, which OpenAI-compatible clients do when you set the token as their API key. Installing Caddy is covered in Caddy as a reverse proxy.

Only change OLLAMA_HOST (for example to 0.0.0.0:11434 or a private interface address) when a client on another host or network really needs direct access, and then restrict port 11434 to known addresses with ufw, for example sudo ufw allow from 10.0.0.0/24 to any port 11434 proto tcp. Make sure ufw is active with only OpenSSH, 80/tcp and 443/tcp open otherwise.

Run Ollama in Docker instead

The official image is ollama/ollama. With Docker installed (Ubuntu, Debian) and, for NVIDIA GPUs, the NVIDIA Container Toolkit set up (Step 2 of How to install vLLM shows the official commands), start it with the API published on loopback only:

Bash
docker run -d --gpus=all -v ollama:/root/.ollama -p 127.0.0.1:11434:11434 --restart unless-stopped --name ollama ollama/ollama
docker exec -it ollama ollama run llama3.2

Leave out --gpus=all for CPU only. For AMD GPUs the documentation uses the ollama/ollama:rocm image with --device /dev/kfd --device /dev/dri instead. Models are stored in the ollama volume. Do not run the container and the systemd service on the same port at the same time.

Back up and restore

Models can always be pulled again, so a backup mainly needs your own work: custom models created with ollama create and their Modelfiles, the systemd override in /etc/systemd/system/ollama.service.d/, and the Ollama home directory with its key pair. Back up the home directory while the service is stopped:

Bash
sudo mkdir -p /opt/backups
sudo systemctl stop ollama
sudo tar czf /opt/backups/ollama-$(date +%F).tar.gz /usr/share/ollama/.ollama /etc/systemd/system/ollama.service.d
sudo systemctl start ollama

If you moved models with OLLAMA_MODELS, add that directory to the archive, or leave it out and pull the models again after a restore. To restore, install Ollama, stop the service, extract the archive from the root directory and fix ownership:

Bash
sudo systemctl stop ollama
sudo tar xzf /opt/backups/ollama-YYYY-MM-DD.tar.gz -C /
sudo chown -R ollama:ollama /usr/share/ollama
sudo systemctl daemon-reload
sudo systemctl start ollama

Copy the archives off the server as well.

Update Ollama

Read the release notes on Ollama's GitHub releases page first. To update, run the install script again, or extract the new tarball over the old installation if you installed manually. The script removes the old libraries before it extracts the new ones:

Bash
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
sh ollama-install.sh
ollama -v

Your models and your systemd override are kept. To install a specific version, for example to roll back, set OLLAMA_VERSION to a version number from the releases page: OLLAMA_VERSION=x.y.z sh ollama-install.sh. For the Docker variant, run docker pull ollama/ollama, remove the container and start it again with the same command.

Troubleshooting

Reading the logs

Ollama logs to the systemd journal. Follow it with journalctl -u ollama --no-pager --follow --pager-end. For more detail, add Environment="OLLAMA_DEBUG=1" with sudo systemctl edit ollama, restart the service and reproduce the problem.

The GPU is not used and ollama ps shows 100% CPU

First confirm the driver with nvidia-smi; if it fails, fix the driver and reboot. If the model is simply larger than the VRAM, Ollama splits it between CPU and GPU or runs it on the CPU; choose a smaller model or tag. After a suspend and resume cycle Ollama can fall back to the CPU; reload the NVIDIA UVM driver with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm. Check the startup lines in the journal for the GPU discovery result, and for AMD GPUs make sure the ollama user is in the render and video groups.

model requires more system memory than is available

The model, its context and any parallel requests do not fit in RAM. Use a smaller model or a more heavily quantized tag, lower OLLAMA_CONTEXT_LENGTH or OLLAMA_NUM_PARALLEL, unload other models with ollama stop, or move to a server with more memory.

403 Forbidden through a reverse proxy or tunnel

Ollama rejects requests whose Host header is not a local address while it listens on 127.0.0.1. Rewrite the header in your proxy, as in the Caddy example in Step 6 (header_up Host localhost:11434).

The ollama command cannot connect to the server

The service is not running or listens somewhere else. Check systemctl status ollama. If you changed OLLAMA_HOST, the CLI needs the same value, for example OLLAMA_HOST=10.0.0.5:11434 ollama list.

Next steps

Frequently asked questions

Does the Ollama API have a password or API key?

No. A local Ollama server accepts every request without authentication, which is why it listens on 127.0.0.1:11434 by default. Keep it there and use an SSH tunnel, a front end such as Open WebUI, or a reverse proxy that checks a token if other machines need access.

Do I need a GPU to run Ollama?

No. Ollama runs models on the CPU when it finds no supported GPU, which is fine for small models and testing. For larger models and faster replies use an NVIDIA GPU with compute capability 5.0 or newer and driver 550 or newer, or an AMD GPU with the ROCm v7 driver.

How much RAM do I need for a model?

Ollama’s model library notes that 7B models generally need at least 8 GB of RAM, 13B models at least 16 GB and 70B models at least 64 GB. Treat the download size on the model page as the floor and add headroom, because longer context and parallel requests use more memory.

Where does Ollama store models and can I move them?

With the Linux installer, models live in /usr/share/ollama/.ollama/models. Set OLLAMA_MODELS in a systemd override to use another directory, and give the ollama user ownership of it.

How do I update Ollama without losing models?

Run the official install script again or extract the new tarball over the old one. Models in the model directory are kept, and the service restarts on the new version.

Sources

Generiraj lozinku

Please confirm