Skip to content

AnythingLLM: your documents, your models, your server

Turn files and websites into workspaces your team can chat with, add agents and MCP tools, and choose the language model per workspace: a local one through Ollama or LocalAI, or a hosted API. Run its Docker edition with multi-user accounts on a HyperDC server, preinstalled as an app option or with our guide.

  • Workspaces that chat with your documents
  • Multi-user accounts and permissions
  • Agents with tools and MCP
  • Runs from 2 GB of RAM and a 2-core CPU
Stack
Node.js, React, Docker image
Default portA new instance has no password: anyone who can reach it can use it and change its settings, so protect it before you publish it.
3001
Minimum
2-core CPU with AVX2, 2 GB RAM, 5 GB disk
Your data
SQLite and LanceDB in the storage folder
Official docs
docs.anythingllm.com

Connects to

  • Ollama
  • LocalAI
  • OpenAI-compatible APIs
  • LanceDB
  • Agents
  • MCP
  • Developer API
  • Embeddable chat

Facts from the project’s official website, documentation and repository, checked in October 2026.

Plans are being prepared

We are preparing ready-to-use plans for AnythingLLM. Tell us how you will use it and how many users you expect, and we will reply with a server that fits. You can also start today on a Linux VPS and install it with our guide.

Which server size fits?

Starting points for vCPU, memory and disk. Grow the server when your data and users grow.

Which server size fits?
Feature
Team with hosted models AnythingLLM calls an API
Recommended Many documents Larger workspaces and more users
Local models With Ollama on a second server
vCPUVirtual processor cores of the server. 2 4 2–4
MemoryMemory for the app, its database and the operating system. 2–4 GB 8 GB 4 GB
DiskUploaded documents, vectors and the database. 20 GB 80 GB 40 GB
Model server Hosted API Hosted API Ollama on a GPU server
Server type Linux VPS VPS or VDS VPS + GPU server
  • Team with hosted models

    AnythingLLM calls an API

    vCPUVirtual processor cores of the server.
    2
    MemoryMemory for the app, its database and the operating system.
    2–4 GB
    DiskUploaded documents, vectors and the database.
    20 GB
    Model server
    Hosted API
    Server type
    Linux VPS
  • Recommended

    Many documents

    Larger workspaces and more users

    vCPUVirtual processor cores of the server.
    4
    MemoryMemory for the app, its database and the operating system.
    8 GB
    DiskUploaded documents, vectors and the database.
    80 GB
    Model server
    Hosted API
    Server type
    VPS or VDS
  • Local models

    With Ollama on a second server

    vCPUVirtual processor cores of the server.
    2–4
    MemoryMemory for the app, its database and the operating system.
    4 GB
    DiskUploaded documents, vectors and the database.
    40 GB
    Model server
    Ollama on a GPU server
    Server type
    VPS + GPU server

AnythingLLM’s documented minimum is 2 GB of RAM, a 2-core CPU with AVX2 and 5 GB of storage. Its documentation recommends running a local model on another machine with a GPU when the AnythingLLM server has none.

Built for working with documents

Each workspace keeps its own documents, model and settings.

Chat with documents

Add PDFs, Word files, websites and more; AnythingLLM embeds them in its built-in LanceDB vector store and answers with citations.

Multi-user mode

The Docker edition supports accounts and permissions for your team; switching to multi-user mode cannot be undone.

Agents and MCP

Let agents browse, summarise and call tools, including MCP servers, inside a workspace.

Developer API

Use workspaces from your own apps through the API documented on your instance at /api/docs.

Your data stays yours

Prompts, files and databases stay on a server you control, in the location you choose, instead of a shared SaaS account.

Full root access

Install what the app needs, change any setting and run more services next to it. Nothing is locked behind a panel.

App option or step-by-step guide

Order the server with the app installed as an option, or set it up yourself on a clean Linux server with our guide.

Near your users

Choose a data center in the United States, Europe or Asia. The order form estimates the latency from where you are to each location.

From order to first login

Order the app preinstalled on your server, or install it yourself with our guide.

  1. Pick the server

    Choose a size from the table above and the data center closest to the people who will use the app.

  2. Add AnythingLLM

    Select AnythingLLM as an app option when you order, or install it on a clean Ubuntu or Debian server with our guide.

  3. Point a domain and enable HTTPS

    Create a DNS record such as app.example.com for the server and put a reverse proxy with a free Let’s Encrypt certificate in front of the app.

  4. Protect it and add your team

    Set a password or switch to multi-user mode before sharing the address, then choose your model and embedder.

Step-by-step setup guides

Install, secure and update the app with our guides, written for current Ubuntu and Debian releases.

More guides

Related solutions

Ollama Hosting

Open LLMs with a private, OpenAI-compatible API

Learn more

Open WebUI Hosting

A private ChatGPT-style chat for your team

Learn more

Dify Hosting

AI apps, agents, RAG and workflows in one studio

Learn more

LLM API Hosting

vLLM, llama.cpp and LocalAI behind an OpenAI API

Learn more

AI & LLM Hosting

Models, chat, agents and AI apps on your servers

Learn more

Frequently asked questions

Still have a question? Send us a message and our team will reply by email.
Is AnythingLLM protected after installing?

No. A new instance has no authentication, so anything that can reach it can use it and change its settings. Set a password for single-user mode or switch to multi-user mode right away, and publish it only through HTTPS.

Does AnythingLLM need a GPU?

No. AnythingLLM itself runs on 2 GB of RAM and a 2-core CPU with AVX2. The language model is the heavy part: use a hosted API, or run Ollama or LocalAI on a separate GPU server, as AnythingLLM’s documentation suggests.

Desktop app or server?

The desktop app is for one person on one computer. The Docker edition on a server adds multi-user accounts and permissions and is reachable for your whole team, which is what we host.

Where are my documents stored?

In the storage folder you mount into the container: the SQLite database, the LanceDB vectors, uploaded files and the .env settings. Back up that folder and never make it world-writable.

Which models can I use?

Local models through Ollama, LocalAI or any generic OpenAI-compatible endpoint, and hosted providers. Each workspace can use its own model, so you can mix fast and accurate ones.

Can I embed the chat on my website?

Yes. AnythingLLM can generate an embeddable chat widget for a workspace, so visitors ask questions about the documents you choose. Limit what the workspace contains, as visitors can see its answers.

Put your documents to work

Tell us the app, how many people will use it and where they are, and we will suggest a server for it.

Şifrə yaradın

Please confirm