How to install Ollama on Ubuntu or Debian and run local LLMs
Install Ollama on Ubuntu or Debian with the official script or tarball, run models on CPU or an NVIDIA GPU, keep the API private and update it safely.
- Beginner
- 30 min read
- Updated
Tested on: Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12, Debian 13
This guide is not available in your language yet, so it is shown in English.
On this page
- Prerequisites
- Step 1 — Prepare the server
- Step 2 — Download, review and run the install script
- Alternative: manual install from the tarball
- Step 3 — Pull and run a model
- Step 4 — Test the REST API
- Step 5 — Configure the service
- Step 6 — Keep the API private and reach it securely
- Run Ollama in Docker instead
- Back up and restore
- Update Ollama
- Troubleshooting
- Reading the logs
- The GPU is not used and ollama ps shows 100% CPU
- model requires more system memory than is available
- 403 Forbidden through a reverse proxy or tunnel
- The ollama command cannot connect to the server
- Next steps
Ollama is a runtime that downloads open large language models (LLMs) and serves them through a simple command line and a REST API on port 11434. It is the easiest way to run models such as Llama, Gemma or Qwen on your own server, and many tools, including Open WebUI, speak to it directly. This guide installs Ollama with its official Linux install script, explains exactly what the script changes, and shows the manual tarball method as an alternative. You then run your first model on CPU or GPU, test the API with curl, move model storage, keep the API private and learn how to update, back up and troubleshoot the service.
Prerequisites
- A server running Ubuntu 24.04 LTS, Ubuntu 26.04 LTS, Debian 12 or Debian 13 on x86_64 (amd64) or arm64.
- A non-root user with
sudorights. If you have not set one up yet, follow Secure a new Linux server and Set up SSH keys first. - Optional, for GPU acceleration: an NVIDIA GPU with compute capability 5.0 or newer and driver 550 or newer (cards with compute capability 5.0 to 6.2 need driver 570 or newer), or a supported AMD Radeon or Instinct GPU with the ROCm v7 driver. See GPU servers for suitable hardware.
- Enough disk space for the models you plan to download.
| Resource | Minimum (official) | Suggested starting point |
|---|---|---|
| CPU | Not published | 4 vCPU for small models on CPU only |
| RAM | Ollama library guidance: 7B models generally need at least 8 GB, 13B at least 16 GB, 70B at least 64 GB | 16 GB for models up to about 8B parameters |
| GPU (optional) | NVIDIA compute capability 5.0+ with driver 550+, or AMD with ROCm v7 | Enough VRAM to hold the whole model, so ollama ps shows 100% GPU |
| Disk | Not published | 50 GB free; each model page lists its download size |
The RAM figures come from Ollama's model library and predate current model families, so treat them as a floor. The suggested values are a conservative starting point, not a benchmark. Memory use grows with the context length and with the number of parallel requests: the documentation states that required memory scales with OLLAMA_NUM_PARALLEL multiplied by OLLAMA_CONTEXT_LENGTH. Without an override, Ollama picks the default context length from the available VRAM: 4k tokens below 24 GiB, 32k tokens from 24 to 48 GiB, and 256k tokens at 48 GiB or more.
Step 1 — Prepare the server
The install script needs curl, zstd (the Linux release archives are .tar.zst files) and lspci or lshw to look for a GPU. Install them:
sudo apt update
sudo apt install curl zstd pciutilsIf the server has an NVIDIA GPU, install the driver before Ollama so you control which driver package is used. The commands below follow NVIDIA's driver installation guide for x86_64 servers: they add NVIDIA's network repository and install the nvidia-open driver. On Debian, NVIDIA's guide also requires the contrib component. The guide enables it with add-apt-repository, which Debian 13 no longer ships, so add contrib next to main in your APT sources yourself (the Components: line in /etc/apt/sources.list.d/debian.sources, or the deb lines in /etc/apt/sources.list, depending on which file your system uses) before you run the Debian commands.
Ubuntu
sudo apt install linux-headers-$(uname -r)
distro=ubuntu$(. /etc/os-release && echo "$VERSION_ID" | tr -d .)
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt install nvidia-open
sudo rebootDebian
sudo apt install linux-headers-$(uname -r)
distro=debian$(. /etc/os-release && echo "$VERSION_ID")
wget https://developer.download.nvidia.com/compute/cuda/repos/$distro/x86_64/cuda-keyring_1.1-1_all.deb
sudo dpkg -i cuda-keyring_1.1-1_all.deb
sudo apt update
sudo apt -V install nvidia-open
sudo rebootAfter the reboot, nvidia-smi should print a table with your GPU, the driver version and the highest CUDA version the driver supports. If you skip this step, the Ollama script detects an NVIDIA GPU without a working driver and installs CUDA drivers itself on Ubuntu and Debian. For AMD GPUs, install the ROCm v7 driver with AMD's amdgpu-install utility as described in AMD's ROCm documentation.
Step 2 — Download, review and run the install script
Ollama's official Linux installer is a shell script. Download it to a file and read it before you run it:
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
less ollama-install.sh
sh ollama-install.shRun the script as your normal sudo user, not with sudo: it calls sudo itself when it needs root, and it adds the user who runs it to the ollama group. The documentation shows the same installer as the one-liner curl -fsSL https://ollama.com/install.sh | sh; downloading first simply lets you review it. The script:
- downloads the Ollama release archive for your CPU architecture and extracts it under
/usr/local(binary in/usr/local/bin, libraries in/usr/local/lib/ollama), removing any older libraries first; - creates a system user and group called
ollamawith the home directory/usr/share/ollama, adds it to therenderandvideogroups when they exist, and adds your user to theollamagroup; - writes
/etc/systemd/system/ollama.service, which runsollama serveas theollamauser withRestart=always, then enables and starts it; - looks for a GPU: it reports an NVIDIA GPU if
nvidia-smiworks, installs CUDA drivers from NVIDIA's repository if it finds an NVIDIA card without a driver, downloads the extra ROCm libraries for AMD cards, and otherwise warns that Ollama will run in CPU-only mode.
When it finishes, it prints that the Ollama API is available at 127.0.0.1:11434. Verify the service and the binary:
systemctl status ollama --no-pager
ollama -v
curl http://127.0.0.1:11434/api/tagsThe service should be active (running), ollama -v prints the version, and the API returns a JSON list of models, which is empty for now.
Alternative: manual install from the tarball
If you prefer not to run the script, extract the release archive yourself. This variant installs into /usr instead of /usr/local. On arm64 servers use ollama-linux-arm64.tar.zst; for AMD GPUs also extract ollama-linux-amd64-rocm.tar.zst the same way.
curl -fsSL https://ollama.com/download/ollama-linux-amd64.tar.zst | sudo tar x -C /usr
sudo useradd -r -s /bin/false -U -m -d /usr/share/ollama ollama
sudo usermod -a -G ollama $(whoami)Then create the service file /etc/systemd/system/ollama.service with this content:
[Unit]
Description=Ollama Service
After=network-online.target
[Service]
ExecStart=/usr/bin/ollama serve
User=ollama
Group=ollama
Restart=always
RestartSec=3
Environment="PATH=/usr/local/sbin:/usr/local/bin:/usr/sbin:/usr/bin:/sbin:/bin"
[Install]
WantedBy=multi-user.targetsudo systemctl daemon-reload
sudo systemctl enable --now ollamaStep 3 — Pull and run a model
Browse the model library at ollama.com/library and check each model's download size. A small model such as llama3.2 (the 3B variant is a 2.0 GB download) is a good first test even without a GPU:
ollama pull llama3.2
ollama run llama3.2ollama run opens an interactive prompt; type a question, and type /bye to leave. You can also pass a prompt directly, for example ollama run llama3.2 "Explain DNS in one sentence". Manage models with these commands:
ollama list
ollama ps
ollama stop llama3.2
ollama rm llama3.2ollama list shows downloaded models, and ollama ps shows loaded models. In the ollama ps output, the PROCESSOR column tells you where the model runs: 100% GPU means it fits completely in VRAM, 100% CPU means it runs from system memory, and a split such as 48%/52% CPU/GPU means it did not fit in VRAM. Models unload after 5 minutes without requests by default.
Step 4 — Test the REST API
The API is what front ends and your own code use. Send a generation request without streaming:
curl http://127.0.0.1:11434/api/generate -d '{
"model": "llama3.2",
"prompt": "Why is the sky blue?",
"stream": false
}'Ollama also offers an OpenAI-compatible API under /v1, so OpenAI client libraries work by pointing their base URL at http://127.0.0.1:11434/v1. The client needs some API key value, which Ollama ignores:
curl http://127.0.0.1:11434/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{"model": "llama3.2", "messages": [{"role": "user", "content": "Say this is a test"}]}'You should see a JSON response with the model's answer in response (native API) or choices[0].message.content (OpenAI API).
Step 5 — Configure the service
Ollama reads its settings from environment variables. Set them for the systemd service with an override file rather than editing the unit, so that updates do not overwrite your changes:
sudo systemctl edit ollamaAdd the variables you need in the [Service] section of the editor that opens, for example:
[Service]
Environment="OLLAMA_CONTEXT_LENGTH=8192"
Environment="OLLAMA_KEEP_ALIVE=10m"sudo systemctl daemon-reload
sudo systemctl restart ollamaThe variables you are most likely to need:
| Variable | Default | What it does |
|---|---|---|
OLLAMA_HOST | 127.0.0.1:11434 | Address and port the API listens on |
OLLAMA_MODELS | /usr/share/ollama/.ollama/models | Model storage directory |
OLLAMA_CONTEXT_LENGTH | Based on VRAM (4k below 24 GiB) | Default context window in tokens |
OLLAMA_KEEP_ALIVE | 5m | How long a model stays loaded after the last request |
OLLAMA_NUM_PARALLEL | 1 | Parallel requests per model |
OLLAMA_MAX_LOADED_MODELS | 3 per GPU, or 3 on CPU | Models kept in memory at the same time |
To keep models on a larger disk, stop the service, copy the existing models, give the ollama user ownership and point OLLAMA_MODELS at the new directory with sudo systemctl edit ollama:
sudo systemctl stop ollama
sudo mkdir -p /srv/ollama/models
sudo cp -a /usr/share/ollama/.ollama/models/. /srv/ollama/models/
sudo chown -R ollama:ollama /srv/ollamaAdd Environment="OLLAMA_MODELS=/srv/ollama/models" to the override, then start Ollama again and confirm with ollama list that your models are still there.
Step 6 — Keep the API private and reach it securely
Ollama has no authentication for local requests. Anyone who can reach port 11434 can run and delete models, so leave OLLAMA_HOST at its default 127.0.0.1:11434 and choose one of these patterns instead of listening on every address:
- From your own computer, open an SSH tunnel and use
http://localhost:11434locally:ssh -L 11434:127.0.0.1:11434 [email protected]. - For a web chat interface, run Open WebUI on the same server. How to run Open WebUI with Ollama shows a setup that reaches Ollama on
127.0.0.1without opening the port. - For remote API clients, put Caddy in front of Ollama and require a token. Generate one with
openssl rand -hex 32, then add a site block:
ollama.example.com {
@authorized header Authorization "Bearer change-me-long-random-token"
handle @authorized {
reverse_proxy 127.0.0.1:11434 {
header_up Host localhost:11434
}
}
handle {
respond 401
}
}Reload Caddy with sudo systemctl reload caddy. The header_up Host line matters: Ollama's documentation rewrites the Host header to localhost:11434 in its proxy examples, and requests with a foreign host name are rejected otherwise. Clients then send Authorization: Bearer with your token, which OpenAI-compatible clients do when you set the token as their API key. Installing Caddy is covered in Caddy as a reverse proxy.
Only change OLLAMA_HOST (for example to 0.0.0.0:11434 or a private interface address) when a client on another host or network really needs direct access, and then restrict port 11434 to known addresses with ufw, for example sudo ufw allow from 10.0.0.0/24 to any port 11434 proto tcp. Make sure ufw is active with only OpenSSH, 80/tcp and 443/tcp open otherwise.
Run Ollama in Docker instead
The official image is ollama/ollama. With Docker installed (Ubuntu, Debian) and, for NVIDIA GPUs, the NVIDIA Container Toolkit set up (Step 2 of How to install vLLM shows the official commands), start it with the API published on loopback only:
docker run -d --gpus=all -v ollama:/root/.ollama -p 127.0.0.1:11434:11434 --restart unless-stopped --name ollama ollama/ollama
docker exec -it ollama ollama run llama3.2Leave out --gpus=all for CPU only. For AMD GPUs the documentation uses the ollama/ollama:rocm image with --device /dev/kfd --device /dev/dri instead. Models are stored in the ollama volume. Do not run the container and the systemd service on the same port at the same time.
Back up and restore
Models can always be pulled again, so a backup mainly needs your own work: custom models created with ollama create and their Modelfiles, the systemd override in /etc/systemd/system/ollama.service.d/, and the Ollama home directory with its key pair. Back up the home directory while the service is stopped:
sudo mkdir -p /opt/backups
sudo systemctl stop ollama
sudo tar czf /opt/backups/ollama-$(date +%F).tar.gz /usr/share/ollama/.ollama /etc/systemd/system/ollama.service.d
sudo systemctl start ollamaIf you moved models with OLLAMA_MODELS, add that directory to the archive, or leave it out and pull the models again after a restore. To restore, install Ollama, stop the service, extract the archive from the root directory and fix ownership:
sudo systemctl stop ollama
sudo tar xzf /opt/backups/ollama-YYYY-MM-DD.tar.gz -C /
sudo chown -R ollama:ollama /usr/share/ollama
sudo systemctl daemon-reload
sudo systemctl start ollamaCopy the archives off the server as well.
Update Ollama
Read the release notes on Ollama's GitHub releases page first. To update, run the install script again, or extract the new tarball over the old installation if you installed manually. The script removes the old libraries before it extracts the new ones:
curl -fsSL https://ollama.com/install.sh -o ollama-install.sh
sh ollama-install.sh
ollama -vYour models and your systemd override are kept. To install a specific version, for example to roll back, set OLLAMA_VERSION to a version number from the releases page: OLLAMA_VERSION=x.y.z sh ollama-install.sh. For the Docker variant, run docker pull ollama/ollama, remove the container and start it again with the same command.
Troubleshooting
Reading the logs
Ollama logs to the systemd journal. Follow it with journalctl -u ollama --no-pager --follow --pager-end. For more detail, add Environment="OLLAMA_DEBUG=1" with sudo systemctl edit ollama, restart the service and reproduce the problem.
The GPU is not used and ollama ps shows 100% CPU
First confirm the driver with nvidia-smi; if it fails, fix the driver and reboot. If the model is simply larger than the VRAM, Ollama splits it between CPU and GPU or runs it on the CPU; choose a smaller model or tag. After a suspend and resume cycle Ollama can fall back to the CPU; reload the NVIDIA UVM driver with sudo rmmod nvidia_uvm && sudo modprobe nvidia_uvm. Check the startup lines in the journal for the GPU discovery result, and for AMD GPUs make sure the ollama user is in the render and video groups.
model requires more system memory than is available
The model, its context and any parallel requests do not fit in RAM. Use a smaller model or a more heavily quantized tag, lower OLLAMA_CONTEXT_LENGTH or OLLAMA_NUM_PARALLEL, unload other models with ollama stop, or move to a server with more memory.
403 Forbidden through a reverse proxy or tunnel
Ollama rejects requests whose Host header is not a local address while it listens on 127.0.0.1. Rewrite the header in your proxy, as in the Caddy example in Step 6 (header_up Host localhost:11434).
The ollama command cannot connect to the server
The service is not running or listens somewhere else. Check systemctl status ollama. If you changed OLLAMA_HOST, the CLI needs the same value, for example OLLAMA_HOST=10.0.0.5:11434 ollama list.
Next steps
- Add a chat interface with Open WebUI and Ollama.
- Serve GGUF models with fine-grained control using llama.cpp server.
- For high-throughput APIs on NVIDIA GPUs, see How to install vLLM.
- Compare servers for local models on the Ollama hosting page.
- Read the official documentation at https://docs.ollama.com for every environment variable.
Frequently asked questions
Does the Ollama API have a password or API key?
No. A local Ollama server accepts every request without authentication, which is why it listens on 127.0.0.1:11434 by default. Keep it there and use an SSH tunnel, a front end such as Open WebUI, or a reverse proxy that checks a token if other machines need access.
Do I need a GPU to run Ollama?
No. Ollama runs models on the CPU when it finds no supported GPU, which is fine for small models and testing. For larger models and faster replies use an NVIDIA GPU with compute capability 5.0 or newer and driver 550 or newer, or an AMD GPU with the ROCm v7 driver.
How much RAM do I need for a model?
Ollama’s model library notes that 7B models generally need at least 8 GB of RAM, 13B models at least 16 GB and 70B models at least 64 GB. Treat the download size on the model page as the floor and add headroom, because longer context and parallel requests use more memory.
Where does Ollama store models and can I move them?
With the Linux installer, models live in /usr/share/ollama/.ollama/models. Set OLLAMA_MODELS in a systemd override to use another directory, and give the ollama user ownership of it.
How do I update Ollama without losing models?
Run the official install script again or extract the new tarball over the old one. Models in the model directory are kept, and the service restarts on the new version.