AnythingLLM: your documents, your models, your server
Turn files and websites into workspaces your team can chat with, add agents and MCP tools, and choose the language model per workspace: a local one through Ollama or LocalAI, or a hosted API. Run its Docker edition with multi-user accounts on a HyperDC server, preinstalled as an app option or with our guide.
- Workspaces that chat with your documents
- Multi-user accounts and permissions
- Agents with tools and MCP
- Runs from 2 GB of RAM and a 2-core CPU
- Stack
- Node.js, React, Docker image
- Default portA new instance has no password: anyone who can reach it can use it and change its settings, so protect it before you publish it.
- 3001
- Minimum
- 2-core CPU with AVX2, 2 GB RAM, 5 GB disk
- Your data
- SQLite and LanceDB in the storage folder
- Official docs
- docs.anythingllm.com
Connects to
- Ollama
- LocalAI
- OpenAI-compatible APIs
- LanceDB
- Agents
- MCP
- Developer API
- Embeddable chat
Facts from the project’s official website, documentation and repository, checked in October 2026.
Plans are being prepared
We are preparing ready-to-use plans for AnythingLLM. Tell us how you will use it and how many users you expect, and we will reply with a server that fits. You can also start today on a Linux VPS and install it with our guide.
Which server size fits?
Starting points for vCPU, memory and disk. Grow the server when your data and users grow.
| Feature |
Team with hosted models
AnythingLLM calls an API
|
Recommended Many documents
Larger workspaces and more users
|
Local models
With Ollama on a second server
|
|---|---|---|---|
| vCPUVirtual processor cores of the server. | 2 | 4 | 2–4 |
| MemoryMemory for the app, its database and the operating system. | 2–4 GB | 8 GB | 4 GB |
| DiskUploaded documents, vectors and the database. | 20 GB | 80 GB | 40 GB |
| Model server | Hosted API | Hosted API | Ollama on a GPU server |
| Server type | Linux VPS | VPS or VDS | VPS + GPU server |
-
Team with hosted models
AnythingLLM calls an API
- vCPUVirtual processor cores of the server.
- 2
- MemoryMemory for the app, its database and the operating system.
- 2–4 GB
- DiskUploaded documents, vectors and the database.
- 20 GB
- Model server
- Hosted API
- Server type
- Linux VPS
-
Recommended
Many documents
Larger workspaces and more users
- vCPUVirtual processor cores of the server.
- 4
- MemoryMemory for the app, its database and the operating system.
- 8 GB
- DiskUploaded documents, vectors and the database.
- 80 GB
- Model server
- Hosted API
- Server type
- VPS or VDS
-
Local models
With Ollama on a second server
- vCPUVirtual processor cores of the server.
- 2–4
- MemoryMemory for the app, its database and the operating system.
- 4 GB
- DiskUploaded documents, vectors and the database.
- 40 GB
- Model server
- Ollama on a GPU server
- Server type
- VPS + GPU server
AnythingLLM’s documented minimum is 2 GB of RAM, a 2-core CPU with AVX2 and 5 GB of storage. Its documentation recommends running a local model on another machine with a GPU when the AnythingLLM server has none.
Built for working with documents
Each workspace keeps its own documents, model and settings.
Chat with documents
Add PDFs, Word files, websites and more; AnythingLLM embeds them in its built-in LanceDB vector store and answers with citations.
Multi-user mode
The Docker edition supports accounts and permissions for your team; switching to multi-user mode cannot be undone.
Agents and MCP
Let agents browse, summarise and call tools, including MCP servers, inside a workspace.
Developer API
Use workspaces from your own apps through the API documented on your instance at /api/docs.
Your data stays yours
Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.
Full root access
Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.
App option or step-by-step guide
Order the server with the app installed as an option, or set it up yourself on a clean Linux server with our guide.
Near your users
Choose a data center in the United States, Europe or Asia. The order form estimates the latency from where you are to each location.
From order to first login
Order the app preinstalled on your server, or install it yourself with our guide.
-
Pick the server
Choose a size from the table above and the data center closest to the people who will use the app.
-
Add AnythingLLM
Select AnythingLLM as an app option when you order, or install it on a clean Ubuntu or Debian server with our guide.
-
Point a domain and enable HTTPS
Create a DNS record such as app.example.com for the server and put a reverse proxy with a free Let’s Encrypt certificate in front of the app.
-
Protect it and add your team
Set a password or switch to multi-user mode before sharing the address, then choose your model and embedder.
Step-by-step setup guides
Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.
-
How to install AnythingLLM with Docker, HTTPS and Ollama
Run the official AnythingLLM Docker image with persistent storage and your own secrets, switch on multi-user mode before going public, publish it over HTTPS with Caddy and connect Ollama.
30 min Intermediate
Related solutions
Frequently asked questions
Is AnythingLLM protected after installing?
No. A new instance has no authentication, so anything that can reach it can use it and change its settings. Set a password for single-user mode or switch to multi-user mode right away, and publish it only through HTTPS.
Does AnythingLLM need a GPU?
No. AnythingLLM itself runs on 2 GB of RAM and a 2-core CPU with AVX2. The language model is the heavy part: use a hosted API, or run Ollama or LocalAI on a separate GPU server, as AnythingLLM’s documentation suggests.
Desktop app or server?
The desktop app is for one person on one computer. The Docker edition on a server adds multi-user accounts and permissions and is reachable for your whole team, which is what we host.
Where are my documents stored?
In the storage folder you mount into the container: the SQLite database, the LanceDB vectors, uploaded files and the .env settings. Back up that folder and never make it world-writable.
Which models can I use?
Local models through Ollama, LocalAI or any generic OpenAI-compatible endpoint, and hosted providers. Each workspace can use its own model, so you can mix fast and accurate ones.
Can I embed the chat on my website?
Yes. AnythingLLM can generate an embeddable chat widget for a workspace, so visitors ask questions about the documents you choose. Limit what the workspace contains, as visitors can see its answers.
Put your documents to work
Tell us the app, how many people will use it and where they are, and we will suggest a server for it.