Ollama on your own server: open models, private API
Pull an open model with one command and serve it to your apps, your team and tools such as Open WebUI, n8n or Dify through an OpenAI-compatible API. Ollama runs on CPU servers for small models or on GPU servers for speed, preinstalled as an app option or set up with our guide.
- Llama, Qwen, Gemma, DeepSeek, Mistral, gpt-oss and Phi
- OpenAI-compatible and Anthropic-compatible API
- CPU servers for small models, GPU servers for speed
- Your prompts never leave your server
- Stack
- Go, Ollama 0.40
- Default portOllama listens on 127.0.0.1 by default and its local API needs no authentication; keep it private and put an authenticated proxy in front if other servers must reach it.
- 11434 (localhost)
- MinimumFrom Ollama’s model pages: at least 16 GB for 13B models and 64 GB for 70B models.
- 7B models: at least 8 GB RAM
- Your dataChange the location with OLLAMA_MODELS, for example to a larger disk.
- Models under /usr/share/ollama/.ollama
- Official docs
- docs.ollama.com
Model families
- Llama
- Qwen
- Gemma
- DeepSeek
- Mistral
- gpt-oss
- Phi
- nomic-embed-text
Facts from the project’s official website, documentation and repository, checked in October 2026.
Plans are being prepared
We are preparing ready-to-use plans for Ollama. Tell us how you will use it and how many users you expect, and we will reply with a server that fits. You can also start today on a Linux VPS and install it with our guide.
Which server size fits?
Starting points for vCPU, memory and disk. Grow the server when your data and users grow.
| Feature |
Small models
7B–8B models on CPU
|
Recommended Mid-size models
13B–14B models, more users
|
Large models
Up to 70B models; GPU for speed
|
|---|---|---|---|
| MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory. | 8–16 GB | 16–32 GB | 64 GB+ |
| vCPUVirtual processor cores of the server. | 4 | 8 | 16+ |
| DiskModel files take several GB each; keep room for the ones you try. | 50 GB | 100 GB | 200 GB+ |
| Models that fit | 7B–8B, quantized | 13B–14B, quantized | Up to 70B |
| Server type | VPS or VDS | VDS or dedicated | Dedicated or GPU server |
-
Small models
7B–8B models on CPU
- MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory.
- 8–16 GB
- vCPUVirtual processor cores of the server.
- 4
- DiskModel files take several GB each; keep room for the ones you try.
- 50 GB
- Models that fit
- 7B–8B, quantized
- Server type
- VPS or VDS
-
Recommended
Mid-size models
13B–14B models, more users
- MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory.
- 16–32 GB
- vCPUVirtual processor cores of the server.
- 8
- DiskModel files take several GB each; keep room for the ones you try.
- 100 GB
- Models that fit
- 13B–14B, quantized
- Server type
- VDS or dedicated
-
Large models
Up to 70B models; GPU for speed
- MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory.
- 64 GB+
- vCPUVirtual processor cores of the server.
- 16+
- DiskModel files take several GB each; keep room for the ones you try.
- 200 GB+
- Models that fit
- Up to 70B
- Server type
- Dedicated or GPU server
From Ollama’s model pages: 7B models generally need at least 8 GB of RAM, 13B models 16 GB and 70B models 64 GB. CPU inference works but is slower; a GPU server gives fast answers. GPU servers are being added to our range step by step; ask us about availability.
What you can do with your own Ollama
One model server for every app and person that needs a language model.
Drop-in API
Point OpenAI or Anthropic client libraries at Ollama’s compatible endpoints, or use the official Python and JavaScript libraries.
Chat with Open WebUI
Add Open WebUI for a ChatGPT-style interface with accounts, history and document search.
Embeddings for RAG
Serve embedding models such as nomic-embed-text or bge-m3 for semantic search in Dify, AnythingLLM or your own code.
CPU or GPU
Ollama runs on CPU, NVIDIA GPUs through CUDA and AMD GPUs through ROCm, and uses the GPU automatically when one is present.
Your data stays yours
Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.
Full root access
Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.
App option or step-by-step guide
Order the server with the app installed as an option, or set it up yourself on a clean Linux server with our guide.
Grow without starting over
Start on a VPS, then move to a bigger plan, a VDS with NVMe storage or a dedicated server when the workload grows.
From order to first login
Order the app preinstalled on your server, or install it yourself with our guide.
-
Pick the server
Choose a size from the table above and the data center closest to the people who will use the app.
-
Add Ollama
Select Ollama as an app option when you order, or install it on a clean Ubuntu or Debian server with our guide.
-
Pull your first model
Run ollama pull with a model from the library, then test it with ollama run or a request to /v1/chat/completions.
-
Connect your apps
Add Open WebUI, n8n or Dify on the same server, or reach Ollama from other servers through an authenticated proxy.
Step-by-step setup guides
Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.
-
How to install Ollama on Ubuntu or Debian and run local LLMs
Install Ollama as a systemd service, pull and run open models on CPU or GPU, test the REST API, move model storage and reach the API securely without exposing port 11434.
30 min Beginner -
How to run Open WebUI with Ollama using Docker Compose and HTTPS
Deploy Open WebUI with Docker Compose next to Ollama, connect them without opening port 11434, create the admin account privately, publish it over HTTPS with Caddy and keep it backed up and updated.
35 min Intermediate
Related solutions
Frequently asked questions
Can Ollama run without a GPU?
Yes. Ollama runs models on the CPU when no GPU is present. Small quantized models answer at a usable speed for one user or background jobs; for larger models or several users at once, choose a GPU server.
How much RAM do I need?
Ollama’s model pages say 7B models generally need at least 8 GB of RAM, 13B models 16 GB and 70B models 64 GB. Leave room for the operating system and the apps next to Ollama, and for longer context windows.
Is the Ollama API protected?
No. The local API on port 11434 needs no authentication, and Ollama binds to 127.0.0.1 by default for that reason. Keep it that way, use it from apps on the same server, or reach it over SSH, a VPN or a reverse proxy that requires a key or a login.
Does Ollama work with OpenAI client libraries?
Yes. Ollama offers OpenAI-compatible endpoints such as /v1/chat/completions, /v1/embeddings and /v1/models, and an Anthropic-compatible API. Point the client’s base URL at your server; the API key value is required by the client but ignored by Ollama.
Which models can I use?
Everything in the Ollama library, including Llama, Qwen, Gemma, DeepSeek, Mistral, gpt-oss and Phi, plus embedding models. You can also import GGUF and Safetensors models.
Where are models stored?
On Linux under /usr/share/ollama/.ollama/models when Ollama runs as a service, or in the volume you mount in Docker. Set OLLAMA_MODELS to keep them on a larger disk.
Why are answers cut short on long documents?
Ollama uses a default context window of 4,096 tokens. Raise the context length in the model settings or the request for long documents; a longer context needs more memory.
How do I update Ollama?
Run the official install script again or pull the new Docker image; your downloaded models stay in place. Update models with ollama pull when a newer version is published.
Run your own models
Tell us the app, how many people will use it and where they are, and we will suggest a server for it.