Run Open WebUI as a production local AI appliance
Timo Wevelsiep•Updated: 15.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Receive Open WebUI ready to use rather than merely installed? The AI Cube Pro combines Open WebUI, a local model, hardening, functional testing, and initial setup into a prepared business AI platform. Explore the AI Cube Pro
Open WebUI is the interface through which staff actually use local AI. It resembles familiar AI chats, can present several models, and manages knowledge, users, groups, and other workspace resources. Production requires the interface to be embedded into a complete platform.
What users receive
- chat with one or more approved models;
- history and workspaces on the organisation's platform;
- file uploads and personal knowledge collections;
- shared team knowledge after approval;
- models, prompts, and tools according to role;
- one interface for local and deliberately approved external models.
These capabilities do not mean an external source is already integrated. Users can maintain knowledge in Open WebUI; automatic Nextcloud, SharePoint, or DMS synchronisation needs a separate pipeline.
The guide to RAG with Nextcloud, SharePoint, and DMS explains how such a connection is built.
The platform underneath
Model endpoint
Ollama, vLLM, or another compatible runtime serves the local model. Model release, quantisation, context, and concurrency are documented. The endpoint is not exposed directly.
Database and storage
Chat history, users, configuration, and metadata need persistent storage. Files and knowledge artefacts use defined volumes or object storage. Backup and recovery must cover both together.
Identity and permissions
Business operation uses local accounts or OIDC/SSO, groups, and minimal defaults. Access to models, knowledge, and tools is deliberate. Administration is separated from normal use.
Network
TLS, internal DNS, firewall, and potentially VPN or NetBird constrain access. Web interface, model server, database, and administration do not need to share one exposed segment.
Minimum standard for production
| Area | Define before release |
|---|---|
| Identity | local accounts or SSO/OIDC or LDAP, MFA at the identity provider, and a managed user lifecycle |
| Authorisation | default role, administrators, groups, and access to models, knowledge, and tools |
| Data | storage locations for chats, uploads, knowledge artefacts, database, logs, and backups |
| Models | approved local and optional external endpoints with unambiguous labels |
| Operations | update path, maintenance window, monitoring, restore test, rollback, and incident responsibility |
| Access | internal use, VPN, NetBird, or reverse proxy with a separate administration path |
These decisions turn an installation into an operable service. They should be documented before several teams store production data on the platform.
Updates without uncontrolled downtime
Open WebUI and the model runtime evolve independently. Updates require release review, backup, test, maintenance window, and rollback. Unreviewed automatic upgrades are unsuitable for a production platform.
Model changes also require evaluation against the same work tasks. A higher version number does not replace quality testing.
Monitoring and backup
Monitor availability, latency, errors, memory, storage, services, and backup status. Content prompts do not need to be sent wholesale to external monitoring. Logs should be limited to what operations require.
A restore test verifies that database, files, keys, and configuration recover together. A backup without a tested restart is not reliable evidence of recoverability.
AI Cube Pro as a prepared foundation
For the AI Cube Pro, WZ-IT installs Open WebUI, local runtime, and an agreed model, configures and hardens the base system, and tests function. Initial setup and five hours of support are included. The platform is usually ready within two weeks after configuration approval.
SSO, automated sources, custom RAG pipelines, MCP servers, special backup targets, and secure site connectivity are added to the existing environment. Optional managed operations start at EUR 149.90 excluding VAT per AI Cube and month.
Sources
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
They use local or approved external models through a familiar chat interface, manage knowledge spaces when permitted, work with files, and use approved models, prompts, and tools.
Open WebUI provides the user and workspace layer. Production also requires model endpoints, database and storage, identity, network, backup, updates, monitoring, and an operating process.
Open WebUI supports groups, permissions, and access to resources such as models and knowledge bases. Externally synchronised sources additionally need source permissions enforced in the RAG pipeline.
Yes. WZ-IT delivers Open WebUI with a local model runtime and agreed model preconfigured. Automated sources and custom workflows are separate integrations.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for confidentiality professions
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





