Application
Installation, configuration, roles, updates and controlled upgrade paths for Open WebUI.
WZ-IT installs, integrates and operates Open WebUI as an AI platform - on our infrastructure or yours. We connect models, GPU resources, knowledge sources, identity and secure network access and extend the solution when needed.
The following are trademarks of their respective owners: Open WebUI (the Open WebUI project). WZ-IT is an independent service provider and has no business, partnership, or contractual relationship with these companies. We offer independent migration, installation, hosting, and operations services.
Open WebUI is an extensible, self-hosted AI interface for Ollama and OpenAI-compatible APIs. It brings models, chats, knowledge sources and tools into one interface and can run locally, in the cloud or as part of a hybrid architecture.
As an independent service provider, we plan more than the application itself: inference, storage, identities, network access, backups, monitoring and integration with existing systems.
We install, integrate and operate Open WebUI on our infrastructure in Germany, in your cloud or on-premises. Local models, external model APIs and hybrid setups are combined according to data classes, performance and budget.
Monitoring, backups, maintenance and human response times follow the agreed operational scope and service level. The current licence protects the Open WebUI branding; white-labelling or modified branding may require a separate enterprise licence from the vendor.
Open WebUI is the visible entry point. Reliable business operations require the application, inference, data, identities, network and operations to be designed together.
Installation, configuration, roles, updates and controlled upgrade paths for Open WebUI.
Ollama, vLLM, LiteLLM or external APIs matched to privacy, latency, load and cost.
Connect RAG, knowledge sources, tools, APIs, SSO and existing business systems in a controlled way.
Monitoring, backups, security updates, incident processes and service levels for the production stack.
The exact scope is defined in the proposal. This matrix shows the typical model for hosting and operations by WZ-IT.
| Area | Responsibility | Typical scope |
|---|---|---|
| Architecture & deployment | WZ-IT | Target architecture, sizing, segmentation, deployment and documented transition into operations. |
| Open WebUI platform | WZ-IT | Installation, technical configuration, controlled updates and recoverability of the application. |
| Monitoring, backup & incidents | WZ-IT | Proactive monitoring, agreed backups and response within the selected service level. |
| Models & providers | Shared | We integrate and operate the technical connection; model selection, usage rights and suitability are agreed with you. |
| Identity & network | Shared | We implement SSO, roles, MFA and secure access while combining your existing identity and network requirements with the target design. |
| Content & AI governance | Customer | You define permitted data, user groups, subject-matter approvals and the responsible use of AI output. |
| Custom extensions | Optional WZ-IT | We deliver pipelines, tools, user interfaces and integrations as clearly scoped additional work. |
Intuitive interface inspired by ChatGPT with full Markdown and LaTeX support for mathematical formulas.
Seamless integration with Ollama for local LLM execution - supports Llama, Mistral, Gemma and many more models.
Local and remote RAG integration - chat with PDFs, Word, Excel, PowerPoint, web pages and YouTube videos.
Create custom models from Ollama base models, add custom characters and agents - directly in the interface.
Open WebUI can run entirely in your infrastructure. External model or search services are connected only when they are part of the agreed architecture.
Granular permissions, user groups and role-based access control for secure team environments.
Modular extension model for custom RAG pipelines, function calling and further integrations, depending on version and configuration.
Use multiple models simultaneously - OpenAI, Anthropic, Ollama and other OpenAI-compatible APIs in parallel.
Integrate web search results directly into RAG conversations with various search providers.
Seamless integration of image generation for dynamic visual content in your chats.
Optimized for desktop, laptop and mobile - installable as Progressive Web App with offline access.
Full internationalization with support for numerous languages - use Open WebUI in your preferred language.
Resource requirements are not determined by Open WebUI alone. Model operations, concurrent users, RAG data, integrations and required resilience are the decisive factors.
| Usage profile | Starting point | What drives sizing |
|---|---|---|
| External model API, small team | CPU S or M | Open WebUI, authentication and the proxy run without local inference. User count, plugins and file processing determine the final requirement. |
| Team platform with RAG | CPU M or L + storage | Document volume, embeddings, vector database, concurrent processing and retention are sized together. |
| Local models or high inference load | GPU assessment | Model size, quantisation, context window, concurrency and target latency determine GPU type, VRAM and GPU count. |
| High availability or multiple sites | Architecture assessment | Load balancing, database, object storage, inference pools, recovery objectives and network paths are designed as one system. |
The pricing calculator covers the standardised CPU workload for the application and operations. GPU resources, model APIs, larger RAG storage, high availability and complex migrations are added after a technical assessment.
The deployment location follows your requirements for control, existing infrastructure, data flows and scaling.
A standardised workload on European infrastructure with operations, monitoring and backup from one provider.
Deployment in your existing cloud account; ownership and the provider contract remain with you while operations can be handled by us.
Operations in your data centre or virtualisation platform, including isolated or connectivity-restricted environments.
The interface, models and knowledge sources can run in different locations connected through controlled network paths.
A clearly defined operating scope instead of an opaque hosting flat fee.
We set up a test instance for you, usually on the next business day. No payment details required. After seven days it is deleted unless you continue.
We combine the right compute size with ongoing operations, backups, monitoring and a service level appropriate for the criticality of Open WebUI. High availability and recovery targets are designed separately where needed.
We also design custom hosting architectures, integrations and migrations around Open WebUI. Contact us for a technical assessment.
One managed standard Open WebUI application is included in the Starter workload. Every service level also includes flexible expert time for planned work during regular service hours. Select compute, additional applications, storage and the appropriate service level.
A workload is one compute instance with the applications agreed for it.
One standard app per workload is already included. Additional dedicated servers count as separate workloads.
€79.90 per started TB and month, including daily encrypted offsite backup with 7-day retention.
Briefly describe the current state and objective for Open WebUI. We assess infrastructure, integration, and ongoing operations.
Professional installation on your infrastructure - on-premise, cloud or hybrid
In your data center
AWS, Azure, Hetzner & more
Advanced architecture after technical and licence assessment
We assess your existing Open WebUI environment, plan the target architecture and safeguard the transition with backups and a rollback option.
Dedicated chat interface with locally controllable data flows; offline operation is possible when models and dependencies remain local
Seamless integration with Ollama for local LLM execution without cloud dependency - Llama, Mistral, Gemma and more
Retrieval Augmented Generation for intelligent document queries and knowledge bases with PDFs, Word, Excel
Local AI models without data transmission to external providers - perfect for regulated industries and data protection
Multi-user environment with roles, permissions and individual chat histories for enterprise teams
Experiment with different LLMs, custom functions, Python tools and API integrations for AI development
We securely connect Open WebUI to users, locations, clouds and existing corporate networks.
NetBird, WireGuard, site connectivity and cloud networking
Keycloak, Authentik and existing directory services
Central rules for users, groups and administrative access
TLS, firewall, rate limiting and traceable access
We connect Open WebUI to users, models, data sources and existing systems through secure network paths. Access, identities and permissions are designed as one architecture.
Full-service installation with no hidden costs
The platform sits between users and approved models, knowledge sources and tools. Access and identity are not treated as afterthoughts.
Employees, teams, external users or distributed offices.
NetBird, WireGuard, reverse proxy, segmentation and controlled exposure.
Keycloak or Authentik, SSO, MFA, groups and role-based permissions.
Shared interface, roles, model access, chats, knowledge and extensions.
Ollama, vLLM, LiteLLM or approved external model APIs.
RAG, vector database, object storage and connected document sources.
APIs, automations, business applications and custom-developed functions.
This is a modular reference design. Required components and their placement are derived from data classes, load, existing IT and operational requirements.
Open WebUI brings chat interface and RAG pipeline - but enterprise requirements demand more. We extend the frontend for your use cases.
We use the REST API for integrations and pipelines for custom model routing: The appropriate model is automatically selected based on input.
Input and output filters: Before sending to the model we can enrich prompts (RAG context), after the response we can perform compliance checks.
For UI customizations we develop custom Svelte components: Special input forms, visualizations, industry-specific chat layouts.
How we implement Open WebUI development in practice.
Different requests (code vs. text vs. analysis) need different models, but users shouldn't have to worry about it.
Pipeline that analyzes input and automatically routes to the optimal model (GPT-4, Claude, local Llama) - transparent to the user.
Internal documents (SharePoint, Confluence) should be searchable in chat without manually uploading them.
Automatic sync of enterprise documents into vector DB. With every question, relevant context from the knowledge base is injected.
LLM responses could contain sensitive info (PII, internal codes). Must filter before display.
Output pipeline that scans responses and masks or blocks sensitive data before displaying to user.
AI Cube Pro combines Open WebUI, local model serving and a model agreed in advance into a fully configured AI platform. Staff can start chatting and create personal or shared knowledge spaces; automated data sources and custom RAG pipelines can be added when required.

Ready for use in under two weeks after configuration approval
EUR 5,999 excl. VAT
RAG, data sources, integrations, secure connectivity and custom workflows are scoped separately.
Optional managed operations from EUR 149.90 excl. VAT per AI Cube and month; §203 scopes include a DPA, limited administrative rights, written secrecy obligations for the people involved and instruction on the criminal consequences
Whether you run Open WebUI in-house or need to host confidentiality-professional data §203-ready - we build, operate and maintain Open WebUI on an encrypted on-site server. Data never leaves the building in cleartext.
See Open WebUI on-premise
Specific answers on hosting, models, RAG, security, operations and licensing.
Yes. The agreed scope can cover the application, operating system and containers, monitoring, backups, updates and incident response. Human response times follow the selected service level; model usage and custom development are priced separately.
Yes. We install and operate Open WebUI in your data centre, on Proxmox, Kubernetes or in your cloud account. Before taking over operations, we assess access, backup, monitoring, technical debt and responsibility boundaries. Learn more about on-premises deployments.
Open WebUI supports Ollama and OpenAI-compatible APIs. This makes it possible to connect local inference systems such as Ollama or vLLM as well as approved external model providers. We select the gateway, models and routing according to data classes, quality, latency and cost.
No. When using external model APIs, the interface normally does not need its own GPU server. Large local language models are sized by model size, VRAM, concurrency and target latency. Where required, we design a managed GPU server.
We design ingestion, text extraction, embeddings, vector storage, metadata and source updates as one pipeline. For confidential content we also address permissions, citations, deletion concepts and the limits of subject-matter reliability.
Yes, provided a technical audit confirms that it can be operated reliably. We review the version, database, volumes, secrets, model connections, backups, network and recovery. Required remediation or migration is scoped transparently before ongoing operations begin.
Standard operations monitor availability and agreed technical states of the workload. Daily backups with seven-day retention are included in the standard offer. Additional databases, vector stores, GPU systems, longer retention and specific RTO/RPO targets are agreed to match the architecture.
For larger environments, we separate the application, database, storage and inference and design load balancing, multiple workloads, inference pools and recovery. High availability is not a generic hosting checkbox; it is designed around failure impact, RTO, RPO and load profile.
The current Open WebUI licence permits use with the original branding left intact. Removing or changing branding is subject to exceptions and may require an enterprise licence; the vendor's current licence is authoritative. We review technical implementation but do not provide legal advice.
Knowledge & guides
Read comparisons, installation guides and architecture articles independently of the specific WZ-IT service.
10.05.2026
Three unauthenticated API calls. No login, no exploit framework, no privilege escalation. Three POST requests to a default port, and the machine's memory is on...
03.12.2025
Open WebUI is a powerful self-hosted web interface for Large Language Models. Whether you want to run local models with Ollama or connect to OpenAI,...
24.11.2025
OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a...
These solutions are often used together with Open WebUI
These solutions offer similar functionalities and can be evaluated together
These solutions are direct alternatives with similar use cases
No risk: worst case, you leave with a clearer understanding of your project than before.


“WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.”
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.