Your own OpenAI-compatible LLM API
Serve open models to your applications through an API your code already speaks. vLLM delivers high throughput on GPUs, llama.cpp runs quantized GGUF models efficiently on CPUs and GPUs, and LocalAI puts many model types behind one endpoint. Run them on HyperDC servers, preinstalled as an app option or with our guide.
- OpenAI-compatible endpoints for chat, completions and embeddings
- vLLM for GPU throughput, llama.cpp for quantized GGUF
- LocalAI for text, images, audio and embeddings
- Keys, proxy and private networking set up safely
- Stack
- Python (vLLM), C/C++ (llama.cpp), Go (LocalAI)
- Default portsllama-server binds to 127.0.0.1 by default; keep every model API private or behind an authenticated proxy.
- vLLM 8000; llama-server and LocalAI 8080
- MinimumvLLM also has CPU builds for x86 and Arm; llama.cpp and LocalAI run on CPU first and use a GPU when present.
- vLLM: GPU with compute capability 7.5+
- Your data
- Model files on local NVMe
- Official docs
- docs.vllm.ai · localai.io
Engines and formats
- vLLM
- llama.cpp
- LocalAI
- GGUF
- Safetensors
- OpenAI API
- Anthropic Messages API
- Prometheus
Facts from the project’s official website, documentation and repository, checked in October 2026.
Plans are being prepared
We are preparing ready-to-use plans for LLM APIs. Tell us how you will use it and how many users you expect, and we will reply with a server that fits. You can also start today on a Linux VPS and install it with our guide.
Which server size fits?
Starting points for vCPU, memory and disk. Grow the server when your data and users grow.
| Feature |
llama.cpp on CPU
Quantized models, light traffic
|
LocalAI
Several model types behind one API
|
Recommended vLLM on GPU
Production traffic, many requests
|
|---|---|---|---|
| vCPUVirtual processor cores of the server. | 4–8 | 8 | Sized to your models |
| MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU. | 16 GB | 16–32 GB | Sized to your models |
| DiskNVMe or SSD storage for the app, its data and local backups. | 50 GB | 100 GB | 200 GB+ |
| Accelerator | CPU | CPU or GPU | GPU |
| Server type | VPS or VDS | VDS or dedicated | GPU server (ask us) |
-
llama.cpp on CPU
Quantized models, light traffic
- vCPUVirtual processor cores of the server.
- 4–8
- MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU.
- 16 GB
- DiskNVMe or SSD storage for the app, its data and local backups.
- 50 GB
- Accelerator
- CPU
- Server type
- VPS or VDS
-
LocalAI
Several model types behind one API
- vCPUVirtual processor cores of the server.
- 8
- MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU.
- 16–32 GB
- DiskNVMe or SSD storage for the app, its data and local backups.
- 100 GB
- Accelerator
- CPU or GPU
- Server type
- VDS or dedicated
-
Recommended
vLLM on GPU
Production traffic, many requests
- vCPUVirtual processor cores of the server.
- Sized to your models
- MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU.
- Sized to your models
- DiskNVMe or SSD storage for the app, its data and local backups.
- 200 GB+
- Accelerator
- GPU
- Server type
- GPU server (ask us)
vLLM requires a GPU with compute capability 7.5 or higher for GPU serving; llama.cpp and LocalAI publish no minimum and run on CPU first. Memory follows the size and quantization of your model: see the Ollama page for typical figures.
Pick the engine for the job
Three maintained open-source servers, one familiar API.
vLLM for throughput
PagedAttention and continuous batching serve many requests at once; FP8, INT4, GPTQ, AWQ and GGUF models are supported.
llama.cpp for efficiency
Plain C/C++ with 1.5- to 8-bit quantization; llama-server adds parallel decoding for several users and a web interface.
LocalAI for everything
Text, embeddings, images and audio from many backends behind one OpenAI-compatible endpoint, with a user and API key system.
Secured the right way
vLLM’s own docs warn not to rely only on --api-key; we keep the engines private and expose them through an authenticated proxy.
Your data stays yours
Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.
Full root access
Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.
App option or step-by-step guide
Order the server with the app installed as an option, or set it up yourself on a clean Linux server with our guide.
Grow without starting over
Start on a VPS, then move to a bigger plan, a VDS with NVMe storage or a dedicated server when the workload grows.
From order to first login
Order the app preinstalled on your server, or install it yourself with our guide.
-
Pick the server
Choose a size from the table above and the data center closest to the people who will use the app.
-
Add vLLM, llama.cpp or LocalAI
Select vLLM, llama.cpp or LocalAI as an app option when you order, or install it on a clean Ubuntu or Debian server with our guide.
-
Load a model
Start vllm serve with a Hugging Face model, llama-server with a GGUF file, or install a model from LocalAI’s gallery.
-
Protect and connect
Set an API key, keep the port private, publish it through a reverse proxy with HTTPS and point your OpenAI client at it.
Step-by-step setup guides
Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.
-
How to run llama.cpp server as an OpenAI-compatible API
Compile llama.cpp, test llama-server with a GGUF model from Hugging Face, run it as a hardened systemd service on 127.0.0.1 with an API key, and publish the OpenAI-compatible API over HTTPS.
40 min Intermediate -
How to install vLLM on an NVIDIA GPU server with Docker
Prepare an NVIDIA GPU server, run vllm/vllm-openai with Docker Compose on 127.0.0.1, secure it with an API key and a Caddy proxy that only exposes /v1, tune memory and multi-GPU settings, or install it with uv and systemd.
45 min Advanced -
How to install LocalAI with Docker Compose as an OpenAI-compatible API
Deploy LocalAI from the official container images, require an API key for every request, install a model from the gallery, test the OpenAI-compatible API and publish it over HTTPS behind Caddy.
35 min Intermediate
Related solutions
Frequently asked questions
vLLM, llama.cpp or LocalAI: which should I use?
Choose vLLM on a GPU server when many requests must be answered quickly. Choose llama.cpp for quantized GGUF models on CPUs or modest GPUs. Choose LocalAI when one endpoint should serve text, embeddings, images and audio from different backends.
Are these APIs compatible with OpenAI clients?
Yes. All three offer OpenAI-compatible endpoints for chat completions, completions, embeddings and models, and vLLM and llama.cpp also support the Anthropic Messages API. Change the client’s base URL and key, and existing code keeps working.
Is an API key enough to secure the server?
Not on its own. vLLM’s --api-key protects only the /v1-style routes and its documentation says not to rely on it alone; llama-server has no key unless you set one, and LocalAI is open until you configure authentication. Keep the engine on a private address and publish it through an authenticated reverse proxy with HTTPS.
Which GPUs does vLLM support?
NVIDIA GPUs with compute capability 7.5 or higher, supported AMD Instinct and Radeon GPUs through ROCm, and Intel data center and Arc GPUs. GPU servers are being added to our range step by step; we check compatibility with your model before you order.
Can I run an LLM API without a GPU?
Yes, with llama.cpp or LocalAI and a quantized model sized to the server’s memory. vLLM also has CPU builds for x86 and Arm. Expect lower throughput than on a GPU.
What is GGUF?
GGUF is the model file format of llama.cpp, with quantized weights from 1.5 to 8 bits per value. Many models are published in GGUF on Hugging Face, and llama-server can download them directly with the -hf option.
Does LocalAI still offer all-in-one images?
No. LocalAI dropped its all-in-one images in version 4.0 to focus on its main images for CPU, NVIDIA CUDA, AMD, Intel and Vulkan. Install the models you need from the gallery instead.
Can I monitor the API?
Yes. vLLM exposes Prometheus metrics, and llama-server does so with the --metrics option. Scrape them with Prometheus and watch latency, tokens per second and queue length.
Serve models to your apps
Tell us the app, how many people will use it and where they are, and we will suggest a server for it.