# Your own OpenAI-compatible LLM API

> Serve open models behind your own OpenAI-compatible API with vLLM, llama.cpp or LocalAI: fast GPU inference or quantized models on CPU, under your control.

Serve open models to your applications through an API your code already speaks. vLLM delivers high throughput on GPUs, llama.cpp runs quantized GGUF models efficiently on CPUs and GPUs, and LocalAI puts many model types behind one endpoint. Run them on HyperDC servers, preinstalled as an app option or with our guide.

- OpenAI-compatible endpoints for chat, completions and embeddings
- vLLM for GPU throughput, llama.cpp for quantized GGUF
- LocalAI for text, images, audio and embeddings
- Keys, proxy and private networking set up safely

See plans Compare server sizes

Stack Python (vLLM), C/C++ (llama.cpp), Go (LocalAI)\
Default ports llama-server binds to 127.0.0.1 by default; keep every model API private or behind an authenticated proxy. vLLM 8000; llama-server and LocalAI 8080\
Minimum vLLM also has CPU builds for x86 and Arm; llama.cpp and LocalAI run on CPU first and use a GPU when present. vLLM: GPU with compute capability 7.5+\
Your data Model files on local NVMe\
Official docs docs.vllm.ai · localai.io

Engines and formats

- vLLM
- llama.cpp
- LocalAI
- GGUF
- Safetensors
- OpenAI API
- Anthropic Messages API
- Prometheus

Facts from the project’s official website, documentation and repository, checked in October 2026.

## Plans are being prepared

We are preparing ready-to-use plans for LLM APIs. Tell us how you will use it and how many users you expect, and we will reply with a server that fits. You can also start today on a Linux VPS and install it with our guide.

[Contact sales](https://hyperdc.com/contact-us) [Open a ticket](https://hyperdc.com/support/new-ticket)

## Which server size fits?

Starting points for vCPU, memory and disk. Grow the server when your data and users grow.

| Feature | **llama.cpp on CPU** Quantized models, light traffic | **LocalAI** Several model types behind one API | Recommended **vLLM on GPU** Production traffic, many requests |
| --- | --- | --- | --- |
| vCPUVirtual processor cores of the server. | 4–8 | 8 | Sized to your models |
| MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU. | 16 GB | 16–32 GB | Sized to your models |
| DiskNVMe or SSD storage for the app, its data and local backups. | 50 GB | 100 GB | 200 GB+ |
| Accelerator | CPU | CPU or GPU | GPU |
| Server type | VPS or VDS | VDS or dedicated | GPU server (ask us) |

- ### llama.cpp on CPU
  Quantized models, light traffic
  vCPUVirtual processor cores of the server. 4–8\
  MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU. 16 GB\
  DiskNVMe or SSD storage for the app, its data and local backups. 50 GB\
  Accelerator CPU\
  Server type VPS or VDS
- ### LocalAI
  Several model types behind one API
  vCPUVirtual processor cores of the server. 8\
  MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU. 16–32 GB\
  DiskNVMe or SSD storage for the app, its data and local backups. 100 GB\
  Accelerator CPU or GPU\
  Server type VDS or dedicated
- Recommended
  ### vLLM on GPU
  Production traffic, many requests
  vCPUVirtual processor cores of the server. Sized to your models\
  MemoryThe quantized model has to fit in RAM on a CPU, or in GPU memory on a GPU. Sized to your models\
  DiskNVMe or SSD storage for the app, its data and local backups. 200 GB+\
  Accelerator GPU\
  Server type GPU server (ask us)

vLLM requires a GPU with compute capability 7.5 or higher for GPU serving; llama.cpp and LocalAI publish no minimum and run on CPU first. Memory follows the size and quantization of your model: see the Ollama page for typical figures.

See plans

## Pick the engine for the job

Three maintained open-source servers, one familiar API.

### vLLM for throughput

PagedAttention and continuous batching serve many requests at once; FP8, INT4, GPTQ, AWQ and GGUF models are supported.

### llama.cpp for efficiency

Plain C/C++ with 1.5- to 8-bit quantization; llama-server adds parallel decoding for several users and a web interface.

### LocalAI for everything

Text, embeddings, images and audio from many backends behind one OpenAI-compatible endpoint, with a user and API key system.

### Secured the right way

vLLM’s own docs warn not to rely only on --api-key; we keep the engines private and expose them through an authenticated proxy.

### Your data stays yours

Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.

### Full root access

Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.

### App option or step-by-step guide

Order the server with the app installed as an option, or set it up yourself on a clean Linux server with our guide.

### Grow without starting over

Start on a VPS, then move to a bigger plan, a VDS with NVMe storage or a dedicated server when the workload grows.

## From order to first login

Order the app preinstalled on your server, or install it yourself with our guide.

[Browse our guides](https://hyperdc.com/guides)

1. ### Pick the server
   Choose a size from the table above and the data center closest to the people who will use the app.
2. ### Add vLLM, llama.cpp or LocalAI
   Select vLLM, llama.cpp or LocalAI as an app option when you order, or install it on a clean Ubuntu or Debian server with our guide.
3. ### Load a model
   Start vllm serve with a Hugging Face model, llama-server with a GGUF file, or install a model from LocalAI’s gallery.
4. ### Protect and connect
   Set an API key, keep the port private, publish it through a reverse proxy with HTTPS and point your OpenAI client at it.

## Step-by-step setup guides

Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.

[More guides](https://hyperdc.com/guides/search?q=llm%20api)

- Tutorials [How to run llama.cpp server as an OpenAI-compatible API](https://hyperdc.com/guides/tutorials/llama-cpp-server)
  Compile llama.cpp, test llama-server with a GGUF model from Hugging Face, run it as a hardened systemd service on 127.0.0.1 with an API key, and publish the OpenAI-compatible API over HTTPS.
  40 min Intermediate
- Tutorials [How to install vLLM on an NVIDIA GPU server with Docker](https://hyperdc.com/guides/tutorials/install-vllm)
  Prepare an NVIDIA GPU server, run vllm/vllm-openai with Docker Compose on 127.0.0.1, secure it with an API key and a Caddy proxy that only exposes /v1, tune memory and multi-GPU settings, or install it with uv and systemd.
  45 min Advanced
- Tutorials [How to install LocalAI with Docker Compose as an OpenAI-compatible API](https://hyperdc.com/guides/tutorials/install-localai)
  Deploy LocalAI from the official container images, require an API key for every request, install a model from the gallery, test the OpenAI-compatible API and publish it over HTTPS behind Caddy.
  35 min Intermediate

## Related solutions

[All solutions](https://hyperdc.com/solutions)

### Ollama Hosting

Open LLMs with a private, OpenAI-compatible API

[Learn more](https://hyperdc.com/ollama-hosting)

### GPU Servers

GPU servers for AI inference and training

[Learn more](https://hyperdc.com/gpu-servers)

### Open WebUI Hosting

A private ChatGPT-style chat for your team

[Learn more](https://hyperdc.com/open-webui-hosting)

### AI Agent Hosting

OpenClaw, Hermes Agent and your own agents, always on

[Learn more](https://hyperdc.com/ai-agents-hosting)

### Dify Hosting

AI apps, agents, RAG and workflows in one studio

[Learn more](https://hyperdc.com/dify-hosting)

### AI & LLM Hosting

Models, chat, agents and AI apps on your servers

[Learn more](https://hyperdc.com/ai-hosting)

## Frequently asked questions

Still have a question? Send us a message and our team will reply by email.

[Contact us](https://hyperdc.com/contact-us) [Open a ticket](https://hyperdc.com/support/new-ticket)

### vLLM, llama.cpp or LocalAI: which should I use?

Choose vLLM on a GPU server when many requests must be answered quickly. Choose llama.cpp for quantized GGUF models on CPUs or modest GPUs. Choose LocalAI when one endpoint should serve text, embeddings, images and audio from different backends.

### Are these APIs compatible with OpenAI clients?

Yes. All three offer OpenAI-compatible endpoints for chat completions, completions, embeddings and models, and vLLM and llama.cpp also support the Anthropic Messages API. Change the client’s base URL and key, and existing code keeps working.

### Is an API key enough to secure the server?

Not on its own. vLLM’s --api-key protects only the /v1-style routes and its documentation says not to rely on it alone; llama-server has no key unless you set one, and LocalAI is open until you configure authentication. Keep the engine on a private address and publish it through an authenticated reverse proxy with HTTPS.

### Which GPUs does vLLM support?

NVIDIA GPUs with compute capability 7.5 or higher, supported AMD Instinct and Radeon GPUs through ROCm, and Intel data center and Arc GPUs. GPU servers are being added to our range step by step; we check compatibility with your model before you order.

### Can I run an LLM API without a GPU?

Yes, with llama.cpp or LocalAI and a quantized model sized to the server’s memory. vLLM also has CPU builds for x86 and Arm. Expect lower throughput than on a GPU.

### What is GGUF?

GGUF is the model file format of llama.cpp, with quantized weights from 1.5 to 8 bits per value. Many models are published in GGUF on Hugging Face, and llama-server can download them directly with the -hf option.

### Does LocalAI still offer all-in-one images?

No. LocalAI dropped its all-in-one images in version 4.0 to focus on its main images for CPU, NVIDIA CUDA, AMD, Intel and Vulkan. Install the models you need from the gallery instead.

### Can I monitor the API?

Yes. vLLM exposes Prometheus metrics, and llama-server does so with the --metrics option. Scrape them with Prometheus and watch latency, tokens per second and queue length.

## Serve models to your apps

Tell us the app, how many people will use it and where they are, and we will suggest a server for it.

See plans [Contact sales](https://hyperdc.com/contact-us)

---

Source: <https://hyperdc.com/llm-api-hosting>\
Updated: 2026-10-09
