Guides about gpu
- How to install Ollama on Ubuntu or Debian and run local LLMs Install Ollama as a systemd service, pull and run open models on CPU or GPU, test the REST API, move model storage and reach the API securely without exposing port 11434. Beginner 30 min
- How to install vLLM on an NVIDIA GPU server with Docker Prepare an NVIDIA GPU server, run vllm/vllm-openai with Docker Compose on 127.0.0.1, secure it with an API key and a Caddy proxy that only exposes /v1, tune memory and multi-GPU settings, or install it with uv and systemd. Advanced 45 min
- How to install LocalAI with Docker Compose as an OpenAI-compatible API Deploy LocalAI from the official container images, require an API key for every request, install a model from the gallery, test the OpenAI-compatible API and publish it over HTTPS behind Caddy. Intermediate 35 min
- How to install ComfyUI on an NVIDIA GPU server securely Set up ComfyUI from the official repository on a GPU server, run it under its own user with systemd, keep it off the public internet, add HTTPS with basic authentication, and manage models, custom nodes, backups and updates. Intermediate 45 min