Open WebUI is a browser-based chat interface for language models you run yourself — a private alternative to ChatGPT’s interface, connected to Ollama on your own server. Your conversations and uploaded documents stay on your VPS. This guide installs it securely, sets realistic expectations for running models on a VPS without a GPU, and covers the security settings that matter most.
What to Expect Without a GPU
Our VPS plans have no GPU, so Ollama runs models on the CPU. That works, with limits:
- Small, quantised models (roughly 1–4 billion parameters) respond at a usable pace for chat, summarising and simple Q&A.
- 7–8 billion parameter models run, but noticeably slower, and need several gigabytes of RAM each.
- Memory decides what you can load: plan for the model’s file size plus a gigabyte or two. More cores make replies faster.
For heavy use or large models, use a hosted model API and keep Open WebUI as the private front end — it can connect to OpenAI-compatible APIs as well as Ollama. VPS for AI applications covers sizing in more detail.
Prerequisites
- Ollama installed and running, with at least one model pulled — see How to Install Ollama and Run Local LLMs on a VPS
- Docker installed
- A domain or subdomain pointed at the server, for HTTPS
Step 1 — Run Open WebUI, Bound to Localhost
docker run -d \
--name open-webui \
--restart unless-stopped \
-p 127.0.0.1:3000:8080 \
--add-host=host.docker.internal:host-gateway \
-e OLLAMA_BASE_URL=http://host.docker.internal:11434 \
-e WEBUI_SECRET_KEY="$(openssl rand -hex 32)" \
-v open-webui:/app/backend/data \
ghcr.io/open-webui/open-webui:main
Note 127.0.0.1:3000. A plain -p 3000:8080 publishes the port to the whole internet, and Docker’s port publishing goes around UFW — your firewall would not stop it. Binding to localhost means the only way in is through Nginx with HTTPS.
Ollama must accept connections from the container. If it only listens on 127.0.0.1, set OLLAMA_HOST=0.0.0.0 in its systemd service and make sure port 11434 is not open in your firewall.
Step 2 — Put Nginx and HTTPS in Front
server {
server_name ai.example.com;
client_max_body_size 50m;
location / {
proxy_pass http://127.0.0.1:3000;
proxy_http_version 1.1;
proxy_set_header Upgrade $http_upgrade;
proxy_set_header Connection "upgrade";
proxy_set_header Host $host;
proxy_set_header X-Forwarded-Proto $scheme;
proxy_read_timeout 300s;
}
}
sudo certbot --nginx -d ai.example.com
The Upgrade headers are needed for streamed replies; the long timeout stops slow CPU responses being cut off. client_max_body_size allows document uploads.
Step 3 — Create the Admin Account, Then Close Sign-Ups
Open the site and create the first account — it becomes the administrator. Then, in Admin Panel → Settings, turn off new sign-ups (or set new users to “pending” so you approve each one). Otherwise anyone who finds the URL can create an account and use your server’s CPU. You can also start the container with -e ENABLE_SIGNUP=false once your accounts exist.
Step 4 — Choose a Model and Chat
Pick an installed model from the dropdown. If none appear, pull one on the host first: ollama pull MODEL_NAME. Start with a small model to see how fast your server is before trying larger ones.
Documents and Knowledge Bases
Open WebUI can index uploaded documents and answer questions about them — a built-in alternative to building your own pipeline (for full control, see How to Deploy a RAG Pipeline on a VPS). Indexing uses CPU and memory too; large document sets on a small server take time.
Several Users
Open WebUI supports multiple accounts with roles and per-model permissions, so a small team can share one server while you keep administrative control.
Updating
docker pull ghcr.io/open-webui/open-webui:main
docker stop open-webui && docker rm open-webui
# then run the same docker run command again; your data stays in the open-webui volume
Backing Up
docker run --rm -v open-webui:/data -v "$PWD":/backup alpine \
tar czf /backup/open-webui-$(date +%F).tar.gz -C /data .
Copy the archive off the server; we do not hold backups of customer servers. Chats, users and uploaded documents all live in that volume.
Common Problems
“Failed to connect to Ollama” — Ollama is not running (systemctl status ollama), or listens only on 127.0.0.1 so the container cannot reach it. Check OLLAMA_HOST and the --add-host flag.
No models in the list — pull at least one model on the host with ollama pull.
Replies stop half-way — the proxy timed out; raise proxy_read_timeout.
The server becomes unresponsive while a model runs — the model is too large for the available RAM. Use a smaller or more heavily quantised model, or a plan with more memory.
Continue Reading
- How to Install Ollama and Run Local LLMs on a VPS
- How to Install Docker Engine on Ubuntu & Debian
- How to Install Let’s Encrypt SSL with Certbot
- VPS for Docker
Browse more articles in AI & Machine Learning on a VPS.