GPU servers for AI inference and training
Run large language models, image generation, fine-tuning and notebooks on GPU servers with full root access, ready for Ollama, vLLM, ComfyUI or your own stack. GPU servers are being added to our range step by step: tell us the models and the number of users, and we will let you know what is available and when.
- Sized for your models and number of users
- Full root access, drivers and Docker GPU support
- Ready for Ollama, vLLM, ComfyUI and Jupyter
- Availability and price confirmed before you order
- GPUsGPU servers are being added step by step; tell us your models and we confirm what is available before you order.
- Ask about availability
- Operating systemsAll four are covered by NVIDIA’s current CUDA toolkit and data center driver guides.
- Ubuntu 24.04/26.04 LTS, Debian 12/13
- ContainersThe NVIDIA Container Toolkit gives Docker, containerd and Podman containers access to the GPU.
- Docker with --gpus all
- Access
- Full root access
- Official docs
- docs.nvidia.com · docs.docker.com
Ready for
- Ollama
- vLLM
- llama.cpp
- ComfyUI
- PyTorch
- Jupyter
- Docker
- NVIDIA Container Toolkit
Facts from the project’s official website, documentation and repository, checked in October 2026.
GPU servers are being added
GPU servers are being added to our range step by step. Tell us the models or workloads you plan to run, how many users you expect and your preferred location, and we will let you know what is available and when.
What to tell us about your workload
A GPU is sized by the model, not by a plan name. These are the questions that decide it.
| Feature |
Recommended Inference
One model for a team
|
Inference at scale
Larger models or many users
|
Training
Fine-tuning and experiments
|
|---|---|---|---|
| GPUs | One GPU | One or more GPUs | Several GPUs |
| GPU memoryThe model, its context and the batch have to fit for full speed. | Enough for the quantized model | Larger models or longer context | Model, gradients and batches |
| System memoryFor the operating system, data loading and the apps next to the model. | Room for the model files | More for parallel requests | More for datasets |
| DiskFast storage for model files, datasets and checkpoints. | NVMe for model files | NVMe for models and logs | NVMe for datasets and checkpoints |
| Typical software | Ollama or llama.cpp | vLLM | PyTorch in Jupyter |
-
Recommended
Inference
One model for a team
- GPUs
- One GPU
- GPU memoryThe model, its context and the batch have to fit for full speed.
- Enough for the quantized model
- System memoryFor the operating system, data loading and the apps next to the model.
- Room for the model files
- DiskFast storage for model files, datasets and checkpoints.
- NVMe for model files
- Typical software
- Ollama or llama.cpp
-
Inference at scale
Larger models or many users
- GPUs
- One or more GPUs
- GPU memoryThe model, its context and the batch have to fit for full speed.
- Larger models or longer context
- System memoryFor the operating system, data loading and the apps next to the model.
- More for parallel requests
- DiskFast storage for model files, datasets and checkpoints.
- NVMe for models and logs
- Typical software
- vLLM
-
Training
Fine-tuning and experiments
- GPUs
- Several GPUs
- GPU memoryThe model, its context and the batch have to fit for full speed.
- Model, gradients and batches
- System memoryFor the operating system, data loading and the apps next to the model.
- More for datasets
- DiskFast storage for model files, datasets and checkpoints.
- NVMe for datasets and checkpoints
- Typical software
- PyTorch in Jupyter
Guidance for your request, not a list of configurations: GPU servers are being added to our range step by step and appear in the plans above as they become available.
What teams run on GPU servers
Workloads where a GPU turns minutes into seconds.
LLM inference
Serve open models to your team or product with Ollama, vLLM or llama.cpp behind an OpenAI-compatible API.
Image and video generation
Run ComfyUI workflows for images and video with the checkpoints and nodes you choose.
Fine-tuning and research
Fine-tune models on your own data with PyTorch in JupyterLab or in containers.
Speech, embeddings and search
Transcribe audio and build embeddings for semantic search and RAG at GPU speed.
Your data stays yours
Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.
Full root access
Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.
Help from our team
Open a ticket from the client area if the server or the network needs attention, and follow every reply in one place.
Android emulators too
Need GPU-accelerated Android emulators instead? See our Android GPU servers.
How to get a GPU server
Tell us the workload; we confirm availability and price before you pay.
-
Tell us your workload
Which models or frameworks, how many users or jobs, and where your users are.
-
Hear what is available
We reply with what we can offer for your workload, the location and the monthly price.
-
Order and connect
After payment we prepare the server and email the login details.
-
Install your stack
Install the driver, CUDA and the NVIDIA Container Toolkit, then Ollama, vLLM or ComfyUI with our guides.
Step-by-step setup guides
Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.
-
How to install Ollama on Ubuntu or Debian and run local LLMs
Install Ollama as a systemd service, pull and run open models on CPU or GPU, test the REST API, move model storage and reach the API securely without exposing port 11434.
30 min Beginner -
How to install vLLM on an NVIDIA GPU server with Docker
Prepare an NVIDIA GPU server, run vllm/vllm-openai with Docker Compose on 127.0.0.1, secure it with an API key and a Caddy proxy that only exposes /v1, tune memory and multi-GPU settings, or install it with uv and systemd.
45 min Advanced -
How to install ComfyUI on an NVIDIA GPU server securely
Set up ComfyUI from the official repository on a GPU server, run it under its own user with systemd, keep it off the public internet, add HTTPS with basic authentication, and manage models, custom nodes, backups and updates.
45 min Intermediate
Related solutions
Frequently asked questions
Which GPU servers do you offer?
GPU servers are being added to our range step by step. Tell us what you want to run and we will tell you what is available, its GPU memory, the location and the price before you order; new GPU plans appear on this page as they go online.
How much GPU memory does my model need?
The model weights, the context and the batch have to fit for full speed. Quantized models need much less memory than full-precision ones, and long contexts need more. Tell us the models and context lengths you plan to use and we will size the server for them.
Which drivers and CUDA versions are supported?
For NVIDIA GPUs, install the data center driver and the CUDA toolkit from NVIDIA’s repositories. NVIDIA’s current guides cover Ubuntu 22.04, 24.04 and 26.04 LTS and Debian 12 and 13, so choose one of these releases.
Can I use the GPU inside Docker?
Yes. With the NVIDIA driver and the NVIDIA Container Toolkit installed, start containers with --gpus all or the matching Compose setting; images such as Ollama’s, vLLM’s and the CUDA builds of Open WebUI then use the GPU.
Do AI tools have GPU requirements of their own?
Yes. Ollama supports NVIDIA GPUs with compute capability 5.0 or higher and driver 550 or newer, as well as AMD GPUs through ROCm. vLLM needs NVIDIA GPUs with compute capability 7.5 or higher, or supported AMD and Intel GPUs. We check this against your models before you order.
Can I run models without a GPU first?
Yes. Small quantized models run on a CPU server with enough memory through Ollama or llama.cpp, which is a good way to start today and test your setup.
Do I get root access?
Yes. You get full root access to install drivers, frameworks and your own software, and our team helps through a support ticket if the server or the network needs attention.
Do you offer GPU servers for Android emulators?
That is a separate line: our Android GPU servers are Windows 11 servers prepared for the LDPlayer emulator; see that page for how they are prepared.
Tell us what you want to run
Send us the models, the number of users and the location, and we will tell you what is available and its price.