Skip to content

Private AI on servers you control

Run open-weight language models, chat interfaces, agents and AI workflows on infrastructure you control. Start on a CPU server for small models and for app builders that call hosted APIs, and plan a GPU server for larger models, image generation and fine-tuning as GPU servers are added to our range.

  • Open models with Ollama, vLLM, llama.cpp or LocalAI
  • Chat, RAG and agents with Open WebUI, Dify and Langflow
  • CPU servers for small models; GPU servers are being added
  • Prompts and documents stay on your server

Pick your AI stack

Model servers, chat interfaces, app builders and agents: each page explains what it needs and how to run it.

Every app runs on a HyperDC VPS, VDS or dedicated server with full root access. Plans appear on each page as soon as they are online.

Private AI for your team and your product

Combine a model server, an interface and your data on servers you control.

A private ChatGPT for your team

Open WebUI with Ollama gives your team a chat with accounts, roles, document search and the models you choose.

An LLM API for your apps

vLLM, llama.cpp and LocalAI serve open models behind an OpenAI-compatible API your code already speaks.

Answers from your documents

Dify, AnythingLLM and Langflow index your files and answer questions with sources.

Agents that work around the clock

OpenClaw, Hermes Agent and n8n agents run scheduled tasks and answer on your chat apps.

Images and media

ComfyUI builds image and video generation workflows on a GPU server.

Notebooks and training

JupyterLab and JupyterHub for experiments, evaluation and fine-tuning next to your data.

Your data stays yours

Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.

Near your users

Choose a data center in the United States, Europe or Asia. The order form estimates the latency from where you are to each location.

Which server for which model?

Memory decides which models fit; a GPU decides how fast they answer.

Which server for which model?
Feature
CPU server Small models and API-based apps
Large-memory server Mid-size models, slower answers
Recommended GPU server Large models, many users, images
Models that fitTypical quantized open models; exact needs depend on the model and its context length. 7B–8B models 13B–14B models Larger models
MemoryMemory for the app, its database and the operating system. 8–16 GB 32–64 GB Sized to your models
vCPUVirtual processor cores of the server. 4 8–16 Sized to your models
Speed Fine for one user and background jobs Usable, slower than a GPU Fast answers for many users
Server type VPS or VDS VDS or dedicated GPU server (ask us)
  • CPU server

    Small models and API-based apps

    Models that fitTypical quantized open models; exact needs depend on the model and its context length.
    7B–8B models
    MemoryMemory for the app, its database and the operating system.
    8–16 GB
    vCPUVirtual processor cores of the server.
    4
    Speed
    Fine for one user and background jobs
    Server type
    VPS or VDS
  • Large-memory server

    Mid-size models, slower answers

    Models that fitTypical quantized open models; exact needs depend on the model and its context length.
    13B–14B models
    MemoryMemory for the app, its database and the operating system.
    32–64 GB
    vCPUVirtual processor cores of the server.
    8–16
    Speed
    Usable, slower than a GPU
    Server type
    VDS or dedicated
  • Recommended

    GPU server

    Large models, many users, images

    Models that fitTypical quantized open models; exact needs depend on the model and its context length.
    Larger models
    MemoryMemory for the app, its database and the operating system.
    Sized to your models
    vCPUVirtual processor cores of the server.
    Sized to your models
    Speed
    Fast answers for many users
    Server type
    GPU server (ask us)

Ollama’s model pages give at least 8 GB of RAM for 7B models, 16 GB for 13B models and 64 GB for 70B models. Quantization lowers the need, longer context raises it. GPU servers are being added to our range step by step; ask us about availability.

From idea to a private AI service

Four steps to your own AI stack.

  1. Choose the stack

    A model server such as Ollama or vLLM, plus an interface or builder such as Open WebUI, Dify or n8n.

  2. Size the server

    Use the table above: memory for the models you want, a GPU when speed or larger models matter.

  3. Install with the app option or our guide

    Order the server with the app preinstalled where the option exists, or follow our step-by-step guides.

  4. Secure and connect

    Keep model APIs private, put interfaces behind HTTPS with accounts, and connect your apps.

Servers and tools around your AI stack

n8n (self-hosted)

Workflow automation and AI agents, self-hosted

Learn more

Docker Hosting

Containers and Compose stacks with full root access

Learn more

Linux VDS Hosting

Dedicated resources and nested virtualization for heavier stacks.

Learn more

Dedicated Servers

A whole machine for large workloads, many apps or GPUs.

Learn more

Android GPU Servers

Android emulators with GPU acceleration on dedicated hardware.

Learn more

Contact us

Questions before you order? Our team will help.

Learn more

Frequently asked questions

Still have a question? Send us a message and our team will reply by email.
Do I need a GPU to run AI models?

Not always. Small quantized models run on a CPU server with enough memory, which is fine for one user, testing and background jobs. A GPU makes answers much faster and is needed for larger models, many users, image generation and training. GPU servers are being added to our range step by step; ask us about availability.

How much memory does a model need?

It depends on the number of parameters, the quantization and the context length. As a guide, Ollama’s model pages ask for at least 8 GB of RAM for 7B models, 16 GB for 13B and 64 GB for 70B models. On a GPU, the model has to fit in GPU memory for full speed.

Which open models can I run?

Any model your model server supports. Ollama’s library includes families such as Llama, Qwen, Gemma, DeepSeek, Mistral, gpt-oss and Phi, plus embedding models for search; vLLM and llama.cpp load models from Hugging Face.

Can I use hosted models like OpenAI or Anthropic as well?

Yes. Open WebUI, Dify, Langflow, n8n and most agent frameworks work with hosted providers and local models side by side, so you can choose per task. Your keys and conversations stay in your server’s database.

Is self-hosted AI more private?

With a local model, prompts and documents never leave your server. When you call a hosted model, the request goes to that provider, but your history, files and settings still stay on your server instead of a shared SaaS account.

How do I keep a model API safe?

Most model servers have no authentication by default: Ollama listens on localhost without a key, and LocalAI and ComfyUI are open until you configure them. Keep them on localhost or a private network, reach them over SSH or a VPN, and publish only interfaces with accounts behind HTTPS.

Do you offer GPU servers?

GPU servers are being added to our range step by step. Tell us the models, the number of users and the location you need, and we will let you know what is available and its price before you order.

Can I start small and upgrade?

Yes. Many teams start with Open WebUI or Dify on a VPS using hosted models, add Ollama with a small model on a larger server, and move inference to a GPU server once usage grows.

Plan your private AI stack

Tell us which models and tools you want to run and for how many users, and we will suggest the servers.

ایجاد گذرواژه

Please confirm