Private AI on servers you control
Run open-weight language models, chat interfaces, agents and AI workflows on infrastructure you control. Start on a CPU server for small models and for app builders that call hosted APIs, and plan a GPU server for larger models, image generation and fine-tuning as GPU servers are added to our range.
- Open models with Ollama, vLLM, llama.cpp or LocalAI
- Chat, RAG and agents with Open WebUI, Dify and Langflow
- CPU servers for small models; GPU servers are being added
- Prompts and documents stay on your server
$ ssh [email protected]
Welcome to Ubuntu 26.04 LTS (GNU/Linux x86_64)
root@vps:~# apt update && apt upgrade -y
0 upgraded, 0 newly installed, 0 to remove.
root@vps:~# ufw allow 443/tcp
Rule added
root@vps:~# systemctl is-active nginx
active
root@vps:~#
Pick your AI stack
Model servers, chat interfaces, app builders and agents: each page explains what it needs and how to run it.
No app matches your search. Tell us what you want to run and we will check it.
Ask about an appEvery app runs on a HyperDC VPS, VDS or dedicated server with full root access. Plans appear on each page as soon as they are online.
Private AI for your team and your product
Combine a model server, an interface and your data on servers you control.
A private ChatGPT for your team
Open WebUI with Ollama gives your team a chat with accounts, roles, document search and the models you choose.
An LLM API for your apps
vLLM, llama.cpp and LocalAI serve open models behind an OpenAI-compatible API your code already speaks.
Answers from your documents
Dify, AnythingLLM and Langflow index your files and answer questions with sources.
Agents that work around the clock
OpenClaw, Hermes Agent and n8n agents run scheduled tasks and answer on your chat apps.
Images and media
ComfyUI builds image and video generation workflows on a GPU server.
Notebooks and training
JupyterLab and JupyterHub for experiments, evaluation and fine-tuning next to your data.
Your data stays yours
Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.
Near your users
Choose a data center in the United States, Europe or Asia. The order form estimates the latency from where you are to each location.
Which server for which model?
Memory decides which models fit; a GPU decides how fast they answer.
| Feature |
CPU server
Small models and API-based apps
|
Large-memory server
Mid-size models, slower answers
|
Recommended GPU server
Large models, many users, images
|
|---|---|---|---|
| Models that fitTypical quantized open models; exact needs depend on the model and its context length. | 7B–8B models | 13B–14B models | Larger models |
| MemoryMemory for the app, its database and the operating system. | 8–16 GB | 32–64 GB | Sized to your models |
| vCPUVirtual processor cores of the server. | 4 | 8–16 | Sized to your models |
| Speed | Fine for one user and background jobs | Usable, slower than a GPU | Fast answers for many users |
| Server type | VPS or VDS | VDS or dedicated | GPU server (ask us) |
-
CPU server
Small models and API-based apps
- Models that fitTypical quantized open models; exact needs depend on the model and its context length.
- 7B–8B models
- MemoryMemory for the app, its database and the operating system.
- 8–16 GB
- vCPUVirtual processor cores of the server.
- 4
- Speed
- Fine for one user and background jobs
- Server type
- VPS or VDS
-
Large-memory server
Mid-size models, slower answers
- Models that fitTypical quantized open models; exact needs depend on the model and its context length.
- 13B–14B models
- MemoryMemory for the app, its database and the operating system.
- 32–64 GB
- vCPUVirtual processor cores of the server.
- 8–16
- Speed
- Usable, slower than a GPU
- Server type
- VDS or dedicated
-
Recommended
GPU server
Large models, many users, images
- Models that fitTypical quantized open models; exact needs depend on the model and its context length.
- Larger models
- MemoryMemory for the app, its database and the operating system.
- Sized to your models
- vCPUVirtual processor cores of the server.
- Sized to your models
- Speed
- Fast answers for many users
- Server type
- GPU server (ask us)
Ollama’s model pages give at least 8 GB of RAM for 7B models, 16 GB for 13B models and 64 GB for 70B models. Quantization lowers the need, longer context raises it. GPU servers are being added to our range step by step; ask us about availability.
From idea to a private AI service
Four steps to your own AI stack.
-
Choose the stack
A model server such as Ollama or vLLM, plus an interface or builder such as Open WebUI, Dify or n8n.
-
Size the server
Use the table above: memory for the models you want, a GPU when speed or larger models matter.
-
Install with the app option or our guide
Order the server with the app preinstalled where the option exists, or follow our step-by-step guides.
-
Secure and connect
Keep model APIs private, put interfaces behind HTTPS with accounts, and connect your apps.
Servers and tools around your AI stack
Frequently asked questions
Do I need a GPU to run AI models?
Not always. Small quantized models run on a CPU server with enough memory, which is fine for one user, testing and background jobs. A GPU makes answers much faster and is needed for larger models, many users, image generation and training. GPU servers are being added to our range step by step; ask us about availability.
How much memory does a model need?
It depends on the number of parameters, the quantization and the context length. As a guide, Ollama’s model pages ask for at least 8 GB of RAM for 7B models, 16 GB for 13B and 64 GB for 70B models. On a GPU, the model has to fit in GPU memory for full speed.
Which open models can I run?
Any model your model server supports. Ollama’s library includes families such as Llama, Qwen, Gemma, DeepSeek, Mistral, gpt-oss and Phi, plus embedding models for search; vLLM and llama.cpp load models from Hugging Face.
Can I use hosted models like OpenAI or Anthropic as well?
Yes. Open WebUI, Dify, Langflow, n8n and most agent frameworks work with hosted providers and local models side by side, so you can choose per task. Your keys and conversations stay in your server’s database.
Is self-hosted AI more private?
With a local model, prompts and documents never leave your server. When you call a hosted model, the request goes to that provider, but your history, files and settings still stay on your server instead of a shared SaaS account.
How do I keep a model API safe?
Most model servers have no authentication by default: Ollama listens on localhost without a key, and LocalAI and ComfyUI are open until you configure them. Keep them on localhost or a private network, reach them over SSH or a VPN, and publish only interfaces with accounts behind HTTPS.
Do you offer GPU servers?
GPU servers are being added to our range step by step. Tell us the models, the number of users and the location you need, and we will let you know what is available and its price before you order.
Can I start small and upgrade?
Yes. Many teams start with Open WebUI or Dify on a VPS using hosted models, add Ollama with a small model on a larger server, and move inference to a GPU server once usage grows.
Plan your private AI stack
Tell us which models and tools you want to run and for how many users, and we will suggest the servers.