Skip to content

Ollama on your own server: open models, private API

Pull an open model with one command and serve it to your apps, your team and tools such as Open WebUI, n8n or Dify through an OpenAI-compatible API. Ollama runs on CPU servers for small models or on GPU servers for speed, preinstalled as an app option or set up with our guide.

  • Llama, Qwen, Gemma, DeepSeek, Mistral, gpt-oss and Phi
  • OpenAI-compatible and Anthropic-compatible API
  • CPU servers for small models, GPU servers for speed
  • Your prompts never leave your server
Stack
Go, Ollama 0.40
Default portOllama listens on 127.0.0.1 by default and its local API needs no authentication; keep it private and put an authenticated proxy in front if other servers must reach it.
11434 (localhost)
MinimumFrom Ollama’s model pages: at least 16 GB for 13B models and 64 GB for 70B models.
7B models: at least 8 GB RAM
Your dataChange the location with OLLAMA_MODELS, for example to a larger disk.
Models under /usr/share/ollama/.ollama
Official docs
docs.ollama.com

Model families

  • Llama
  • Qwen
  • Gemma
  • DeepSeek
  • Mistral
  • gpt-oss
  • Phi
  • nomic-embed-text

Facts from the project’s official website, documentation and repository, checked in October 2026.

Plans are being prepared

We are preparing ready-to-use plans for Ollama. Tell us how you will use it and how many users you expect, and we will reply with a server that fits. You can also start today on a Linux VPS and install it with our guide.

Which server size fits?

Starting points for vCPU, memory and disk. Grow the server when your data and users grow.

Which server size fits?
Feature
Small models 7B–8B models on CPU
Recommended Mid-size models 13B–14B models, more users
Large models Up to 70B models; GPU for speed
MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory. 8–16 GB 16–32 GB 64 GB+
vCPUVirtual processor cores of the server. 4 8 16+
DiskModel files take several GB each; keep room for the ones you try. 50 GB 100 GB 200 GB+
Models that fit 7B–8B, quantized 13B–14B, quantized Up to 70B
Server type VPS or VDS VDS or dedicated Dedicated or GPU server
  • Small models

    7B–8B models on CPU

    MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory.
    8–16 GB
    vCPUVirtual processor cores of the server.
    4
    DiskModel files take several GB each; keep room for the ones you try.
    50 GB
    Models that fit
    7B–8B, quantized
    Server type
    VPS or VDS
  • Recommended

    Mid-size models

    13B–14B models, more users

    MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory.
    16–32 GB
    vCPUVirtual processor cores of the server.
    8
    DiskModel files take several GB each; keep room for the ones you try.
    100 GB
    Models that fit
    13B–14B, quantized
    Server type
    VDS or dedicated
  • Large models

    Up to 70B models; GPU for speed

    MemoryRAM for CPU inference; on a GPU the model should fit in GPU memory.
    64 GB+
    vCPUVirtual processor cores of the server.
    16+
    DiskModel files take several GB each; keep room for the ones you try.
    200 GB+
    Models that fit
    Up to 70B
    Server type
    Dedicated or GPU server

From Ollama’s model pages: 7B models generally need at least 8 GB of RAM, 13B models 16 GB and 70B models 64 GB. CPU inference works but is slower; a GPU server gives fast answers. GPU servers are being added to our range step by step; ask us about availability.

What you can do with your own Ollama

One model server for every app and person that needs a language model.

Drop-in API

Point OpenAI or Anthropic client libraries at Ollama’s compatible endpoints, or use the official Python and JavaScript libraries.

Chat with Open WebUI

Add Open WebUI for a ChatGPT-style interface with accounts, history and document search.

Embeddings for RAG

Serve embedding models such as nomic-embed-text or bge-m3 for semantic search in Dify, AnythingLLM or your own code.

CPU or GPU

Ollama runs on CPU, NVIDIA GPUs through CUDA and AMD GPUs through ROCm, and uses the GPU automatically when one is present.

Your data stays yours

Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.

Full root access

Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.

App option or step-by-step guide

Order the server with the app installed as an option, or set it up yourself on a clean Linux server with our guide.

Grow without starting over

Start on a VPS, then move to a bigger plan, a VDS with NVMe storage or a dedicated server when the workload grows.

From order to first login

Order the app preinstalled on your server, or install it yourself with our guide.

  1. Pick the server

    Choose a size from the table above and the data center closest to the people who will use the app.

  2. Add Ollama

    Select Ollama as an app option when you order, or install it on a clean Ubuntu or Debian server with our guide.

  3. Pull your first model

    Run ollama pull with a model from the library, then test it with ollama run or a request to /v1/chat/completions.

  4. Connect your apps

    Add Open WebUI, n8n or Dify on the same server, or reach Ollama from other servers through an authenticated proxy.

Step-by-step setup guides

Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.

More guides

Related solutions

Open WebUI Hosting

A private ChatGPT-style chat for your team

Learn more

GPU Servers

GPU servers for AI inference and training

Learn more

LLM API Hosting

vLLM, llama.cpp and LocalAI behind an OpenAI API

Learn more

AnythingLLM Hosting

Chat with your documents, with agents and MCP

Learn more

n8n (self-hosted)

Workflow automation and AI agents, self-hosted

Learn more

AI & LLM Hosting

Models, chat, agents and AI apps on your servers

Learn more

Frequently asked questions

Still have a question? Send us a message and our team will reply by email.
Can Ollama run without a GPU?

Yes. Ollama runs models on the CPU when no GPU is present. Small quantized models answer at a usable speed for one user or background jobs; for larger models or several users at once, choose a GPU server.

How much RAM do I need?

Ollama’s model pages say 7B models generally need at least 8 GB of RAM, 13B models 16 GB and 70B models 64 GB. Leave room for the operating system and the apps next to Ollama, and for longer context windows.

Is the Ollama API protected?

No. The local API on port 11434 needs no authentication, and Ollama binds to 127.0.0.1 by default for that reason. Keep it that way, use it from apps on the same server, or reach it over SSH, a VPN or a reverse proxy that requires a key or a login.

Does Ollama work with OpenAI client libraries?

Yes. Ollama offers OpenAI-compatible endpoints such as /v1/chat/completions, /v1/embeddings and /v1/models, and an Anthropic-compatible API. Point the client’s base URL at your server; the API key value is required by the client but ignored by Ollama.

Which models can I use?

Everything in the Ollama library, including Llama, Qwen, Gemma, DeepSeek, Mistral, gpt-oss and Phi, plus embedding models. You can also import GGUF and Safetensors models.

Where are models stored?

On Linux under /usr/share/ollama/.ollama/models when Ollama runs as a service, or in the volume you mount in Docker. Set OLLAMA_MODELS to keep them on a larger disk.

Why are answers cut short on long documents?

Ollama uses a default context window of 4,096 tokens. Raise the context length in the model settings or the request for long documents; a longer context needs more memory.

How do I update Ollama?

Run the official install script again or pull the new Docker image; your downloaded models stay in place. Update models with ollama pull when a newer version is published.

Run your own models

Tell us the app, how many people will use it and where they are, and we will suggest a server for it.

Jelszó létrehozása

Please confirm