Guides about openai api
- How to run llama.cpp server as an OpenAI-compatible API Compile llama.cpp, test llama-server with a GGUF model from Hugging Face, run it as a hardened systemd service on 127.0.0.1 with an API key, and publish the OpenAI-compatible API over HTTPS. Intermediate 40 min
- How to install vLLM on an NVIDIA GPU server with Docker Prepare an NVIDIA GPU server, run vllm/vllm-openai with Docker Compose on 127.0.0.1, secure it with an API key and a Caddy proxy that only exposes /v1, tune memory and multi-GPU settings, or install it with uv and systemd. Advanced 45 min
- How to install LocalAI with Docker Compose as an OpenAI-compatible API Deploy LocalAI from the official container images, require an API key for every request, install a model from the gallery, test the OpenAI-compatible API and publish it over HTTPS behind Caddy. Intermediate 35 min