Skip to content

Your own OpenAI-compatible LLM API

Serve open models to your applications through an API your code already speaks. vLLM delivers high throughput on GPUs, llama.cpp runs quantized GGUF models efficiently on CPUs and GPUs, and LocalAI puts many model types behind one endpoint. Run them on HyperDC servers, preinstalled as an app option or with our guide.

  • OpenAI-compatible endpoints for chat, completions and embeddings
  • vLLM for GPU throughput, llama.cpp for quantized GGUF
  • LocalAI for text, images, audio and embeddings
  • Keys, proxy and private networking set up safely
Stack
Python (vLLM), C/C++ (llama.cpp), Go (LocalAI)
Default portsllama-server binds to 127.0.0.1 by default; keep every model API private or behind an authenticated proxy.
vLLM 8000; llama-server and LocalAI 8080
MinimumvLLM also has CPU builds for x86 and Arm; llama.cpp and LocalAI run on CPU first and use a GPU when present.
vLLM: GPU with compute capability 7.5+
Your data
Model files on local NVMe
Official docs
docs.vllm.ai · localai.io

Engines and formats

  • vLLM
  • llama.cpp
  • LocalAI
  • GGUF
  • Safetensors
  • OpenAI API
  • Anthropic Messages API
  • Prometheus

Facts from the project’s official website, documentation and repository, checked in October 2026.

Plans are being prepared

We are preparing ready-to-use plans for LLM APIs. Tell us how you will use it and how many users you expect, and we will reply with a server that fits. You can also start today on a Linux VPS and install it with our guide.

Which server size fits?

Starting points for vCPU, memory and disk. Grow the server when your data and users grow.

Which server size fits?
Feature
llama.cpp on CPU Quantized models, light traffic
LocalAI Several model types behind one API
Recommended vLLM on GPU Production traffic, many requests
vCPUVirtual processor cores of the server. 4–8 8 Sized to your models
MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU. 16 GB 16–32 GB Sized to your models
DiskNVMe or SSD storage for the app, its data and local backups. 50 GB 100 GB 200 GB+
Accelerator CPU CPU or GPU GPU
Server type VPS or VDS VDS or dedicated GPU server (ask us)
  • llama.cpp on CPU

    Quantized models, light traffic

    vCPUVirtual processor cores of the server.
    4–8
    MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU.
    16 GB
    DiskNVMe or SSD storage for the app, its data and local backups.
    50 GB
    Accelerator
    CPU
    Server type
    VPS or VDS
  • LocalAI

    Several model types behind one API

    vCPUVirtual processor cores of the server.
    8
    MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU.
    16–32 GB
    DiskNVMe or SSD storage for the app, its data and local backups.
    100 GB
    Accelerator
    CPU or GPU
    Server type
    VDS or dedicated
  • Recommended

    vLLM on GPU

    Production traffic, many requests

    vCPUVirtual processor cores of the server.
    Sized to your models
    MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU.
    Sized to your models
    DiskNVMe or SSD storage for the app, its data and local backups.
    200 GB+
    Accelerator
    GPU
    Server type
    GPU server (ask us)

vLLM requires a GPU with compute capability 7.5 or higher for GPU serving; llama.cpp and LocalAI publish no minimum and run on CPU first. Memory follows the size and quantization of your model: see the Ollama page for typical figures.

Pick the engine for the job

Three maintained open-source servers, one familiar API.

vLLM for throughput

PagedAttention and continuous batching serve many requests at once; FP8, INT4, GPTQ, AWQ and GGUF models are supported.

llama.cpp for efficiency

Plain C/C++ with 1.5- to 8-bit quantization; llama-server adds parallel decoding for several users and a web interface.

LocalAI for everything

Text, embeddings, images and audio from many backends behind one OpenAI-compatible endpoint, with a user and API key system.

Secured the right way

vLLM’s own docs warn not to rely only on --api-key; we keep the engines private and expose them through an authenticated proxy.

Your data stays yours

Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.

Full root access

Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.

App option or step-by-step guide

Order the server with the app installed as an option, or set it up yourself on a clean Linux server with our guide.

Grow without starting over

Start on a VPS, then move to a bigger plan, a VDS with NVMe storage or a dedicated server when the workload grows.

From order to first login

Order the app preinstalled on your server, or install it yourself with our guide.

  1. Pick the server

    Choose a size from the table above and the data center closest to the people who will use the app.

  2. Add vLLM, llama.cpp or LocalAI

    Select vLLM, llama.cpp or LocalAI as an app option when you order, or install it on a clean Ubuntu or Debian server with our guide.

  3. Load a model

    Start vllm serve with a Hugging Face model, llama-server with a GGUF file, or install a model from LocalAI’s gallery.

  4. Protect and connect

    Set an API key, keep the port private, publish it through a reverse proxy with HTTPS and point your OpenAI client at it.

Step-by-step setup guides

Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.

More guides

Related solutions

Ollama Hosting

Open LLMs with a private, OpenAI-compatible API

Learn more

GPU Servers

GPU servers for AI inference and training

Learn more

Open WebUI Hosting

A private ChatGPT-style chat for your team

Learn more

AI Agent Hosting

OpenClaw, Hermes Agent and your own agents, always on

Learn more

Dify Hosting

AI apps, agents, RAG and workflows in one studio

Learn more

AI & LLM Hosting

Models, chat, agents and AI apps on your servers

Learn more

Frequently asked questions

Still have a question? Send us a message and our team will reply by email.
vLLM, llama.cpp or LocalAI: which should I use?

Choose vLLM on a GPU server when many requests must be answered quickly. Choose llama.cpp for quantized GGUF models on CPUs or modest GPUs. Choose LocalAI when one endpoint should serve text, embeddings, images and audio from different backends.

Are these APIs compatible with OpenAI clients?

Yes. All three offer OpenAI-compatible endpoints for chat completions, completions, embeddings and models, and vLLM and llama.cpp also support the Anthropic Messages API. Change the client’s base URL and key, and existing code keeps working.

Is an API key enough to secure the server?

Not on its own. vLLM’s --api-key protects only the /v1-style routes and its documentation says not to rely on it alone; llama-server has no key unless you set one, and LocalAI is open until you configure authentication. Keep the engine on a private address and publish it through an authenticated reverse proxy with HTTPS.

Which GPUs does vLLM support?

NVIDIA GPUs with compute capability 7.5 or higher, supported AMD Instinct and Radeon GPUs through ROCm, and Intel data center and Arc GPUs. GPU servers are being added to our range step by step; we check compatibility with your model before you order.

Can I run an LLM API without a GPU?

Yes, with llama.cpp or LocalAI and a quantized model sized to the server’s memory. vLLM also has CPU builds for x86 and Arm. Expect lower throughput than on a GPU.

What is GGUF?

GGUF is the model file format of llama.cpp, with quantized weights from 1.5 to 8 bits per value. Many models are published in GGUF on Hugging Face, and llama-server can download them directly with the -hf option.

Does LocalAI still offer all-in-one images?

No. LocalAI dropped its all-in-one images in version 4.0 to focus on its main images for CPU, NVIDIA CUDA, AMD, Intel and Vulkan. Install the models you need from the gallery instead.

Can I monitor the API?

Yes. vLLM exposes Prometheus metrics, and llama-server does so with the --metrics option. Scrape them with Prometheus and watch latency, tokens per second and queue length.

Serve models to your apps

Tell us the app, how many people will use it and where they are, and we will suggest a server for it.

產生密碼

Please confirm