Application
Installation, configuration, roles, updates and controlled upgrade paths for Open WebUI.
WZ-IT installs, integrates and operates Open WebUI as an AI platform - on our infrastructure or yours. We connect models, GPU resources, knowledge sources, identity and secure network access and extend the solution when needed.
Companies worldwide trust WZ-IT
The following are trademarks of their respective owners: Open WebUI (the Open WebUI project). WZ-IT is an independent service provider and has no business, partnership, or contractual relationship with these companies. We offer independent migration, installation, hosting, and operations services.
Open WebUI is an extensible, self-hosted AI interface for Ollama and OpenAI-compatible APIs. It brings models, chats, knowledge sources and tools into one interface and can run locally, in the cloud or as part of a hybrid architecture.
As an independent service provider, we plan more than the application itself: inference, storage, identities, network access, backups, monitoring and integration with existing systems.
We install, integrate and operate Open WebUI on our infrastructure in Germany, in your cloud or on-premises. Local models, external model APIs and hybrid setups are combined according to data classes, performance and budget.
Monitoring, backups, maintenance and human response times follow the agreed operational scope and service level. The current licence protects the Open WebUI branding; white-labelling or modified branding may require a separate enterprise licence from the vendor.
AI Cube combines Open WebUI, local model serving and a model agreed in advance into a fully configured AI platform. Staff can start chatting and create personal or shared knowledge spaces; automated data sources and custom RAG pipelines can be added when required.

Delivered ready to use within 10 working days
€6,490 excl. VAT one-time
Device, setup, briefing and shipping
€349.90 excl. VAT / month
AI Cube Care · first year accompanied, then cancellable monthly
RAG, data sources, integrations, secure connectivity and custom workflows are scoped separately.
AI Cube Care keeps the device running: a monthly maintenance window, tested model updates with a rollback path, round-the-clock monitoring, ticket support and warranty handling. For confidentiality professionals we add a data processing agreement under Art. 28 GDPR and the written secrecy obligation under Section 203 (4) of the German Criminal Code, including instruction on the criminal consequences; on request remote access stays switched off by default.
Open WebUI is the visible entry point. Reliable business operations require the application, inference, data, identities, network and operations to be designed together.
Installation, configuration, roles, updates and controlled upgrade paths for Open WebUI.
Ollama, vLLM, LiteLLM or external APIs matched to privacy, latency, load and cost.
Connect RAG, knowledge sources, tools, APIs, SSO and existing business systems in a controlled way.
Monitoring, backups, security updates, incident processes and service levels for the production stack.
The exact scope is defined in the proposal. This matrix shows the typical model for hosting and operations by WZ-IT.
| Area | Responsibility | Typical scope |
|---|---|---|
| Architecture & deployment | WZ-IT | Target architecture, sizing, segmentation, deployment and documented transition into operations. |
| Open WebUI platform | WZ-IT | Installation, technical configuration, controlled updates and recoverability of the application. |
| Monitoring, backup & incidents | WZ-IT | Proactive monitoring, agreed backups and response within the selected service level. |
| Models & providers | Shared | We integrate and operate the technical connection; model selection, usage rights and suitability are agreed with you. |
| Identity & network | Shared | We implement SSO, roles, MFA and secure access while combining your existing identity and network requirements with the target design. |
| Content & AI governance | Customer | You define permitted data, user groups, subject-matter approvals and the responsible use of AI output. |
| Custom extensions | Optional WZ-IT | We deliver pipelines, tools, user interfaces and integrations as clearly scoped additional work. |
Intuitive interface inspired by ChatGPT with full Markdown and LaTeX support for mathematical formulas.
Seamless integration with Ollama for local LLM execution - supports Llama, Mistral, Gemma and many more models.
Local and remote RAG integration - chat with PDFs, Word, Excel, PowerPoint, web pages and YouTube videos.
Create custom models from Ollama base models, add custom characters and agents - directly in the interface.
Open WebUI can run entirely in your infrastructure. External model or search services are connected only when they are part of the agreed architecture.
Granular permissions, user groups and role-based access control for secure team environments.
Modular extension model for custom RAG pipelines, function calling and further integrations, depending on version and configuration.
Use multiple models simultaneously - OpenAI, Anthropic, Ollama and other OpenAI-compatible APIs in parallel.
Integrate web search results directly into RAG conversations with various search providers.
Seamless integration of image generation for dynamic visual content in your chats.
Optimized for desktop, laptop and mobile - installable as Progressive Web App with offline access.
Full internationalization with support for numerous languages - use Open WebUI in your preferred language.
Resource requirements are not determined by Open WebUI alone. Model operations, concurrent users, RAG data, integrations and required resilience are the decisive factors.
| Usage profile | Starting point | What drives sizing |
|---|---|---|
| External model API, small team | CPU S or M | Open WebUI, authentication and the proxy run without local inference. User count, plugins and file processing determine the final requirement. |
| Team platform with RAG | CPU M or L + storage | Document volume, embeddings, vector database, concurrent processing and retention are sized together. |
| Local models or high inference load | GPU assessment | Model size, quantisation, context window, concurrency and target latency determine GPU type, VRAM and GPU count. |
| High availability or multiple sites | Architecture assessment | Load balancing, database, object storage, inference pools, recovery objectives and network paths are designed as one system. |
The pricing calculator covers the standardised CPU workload for the application and operations. GPU resources, model APIs, larger RAG storage, high availability and complex migrations are added after a technical assessment.
The deployment location follows your requirements for control, existing infrastructure, data flows and scaling.
A standardised workload on European infrastructure with operations, monitoring and backup from one provider.
Deployment in your existing cloud account; ownership and the provider contract remain with you while operations can be handled by us.
Operations in your data centre or virtualisation platform, including isolated or connectivity-restricted environments.
The interface, models and knowledge sources can run in different locations connected through controlled network paths.
A clearly defined operating scope instead of an opaque hosting flat fee.
We set up a test instance for you, usually on the next business day. No payment details required. After seven days it is deleted unless you continue.
We combine the right compute size with ongoing operations, backups, monitoring and a service level appropriate for the criticality of Open WebUI. High availability and recovery targets are designed separately where needed.
We also design custom hosting architectures, integrations and migrations around Open WebUI. Contact us for a technical assessment.
One managed standard Open WebUI application is included in the Starter workload. Business and higher levels add a flexible operations allowance for planned work during regular service hours. Select compute, additional applications, storage and the appropriate service level.
A workload is one compute instance with the applications agreed for it.
One standard app per workload is already included. Additional dedicated servers count as separate workloads.
€7.99 excl. VAT per additional 100 GB per month, including daily encrypted offsite backup with 7-day retention. Billing is based on provisioned capacity, not actual usage.
We also implement custom backup schedules and retention periods. For example: daily backups retained for 30 days and one monthly backup retained for 12 months. Scope, storage requirements and additional costs are agreed separately.
+€299 excluding VAT per workload and month
Standard hosting runs on a single instance without high availability. Availability for custom architectures is agreed separately.
Zero downtime on request
Distributed operation with several instances, a database cluster and rolling updates, so that a node failure causes no interruption. Only for applications suited to it; design and price after technical assessment.
This page covers installation and the agreed platform operations for Open WebUI. Custom code, the access layer and special confidentiality requirements remain clearly separated responsibilities that can be added when needed.
Maintain the application
Keep custom extensions, interfaces and dependencies controlled, updated and supportable.
from €699.90 excl. VAT / month
View serviceSecure access
Add identity-based access and private network paths as a separate operational layer.
from €349.90 excl. VAT / month
View serviceOperate confidentially
For professional secrecy holders, define contractual, access and infrastructure requirements in a dedicated operating model.
from €499.90 excl. VAT / month
View serviceThree relevant adjacent paths instead of a long list of further products.
All prices are net and exclude statutory VAT. The offers are addressed to businesses.
View the overall systemEnquiry
Briefly describe the current state and objective for Open WebUI. We assess infrastructure, integration, and ongoing operations.
Professional installation on your infrastructure - on-premise, cloud or hybrid
In your data center
AWS, Azure, Hetzner & more
Advanced architecture after technical and licence assessment
We assess your existing Open WebUI environment, plan the target architecture and safeguard the transition with backups and a rollback option.
Dedicated chat interface with locally controllable data flows; offline operation is possible when models and dependencies remain local
Seamless integration with Ollama for local LLM execution without cloud dependency - Llama, Mistral, Gemma and more
Retrieval Augmented Generation for intelligent document queries and knowledge bases with PDFs, Word, Excel
Local AI models without data transmission to external providers - perfect for regulated industries and data protection
Multi-user environment with roles, permissions and individual chat histories for enterprise teams
Experiment with different LLMs, custom functions, Python tools and API integrations for AI development
We securely connect Open WebUI to users, locations, clouds and existing corporate networks.
NetBird, WireGuard, site connectivity and cloud networking
Keycloak, Authentik and existing directory services
Central rules for users, groups and administrative access
TLS, firewall, rate limiting and traceable access
We connect Open WebUI to users, models, data sources and existing systems through secure network paths. Access, identities and permissions are designed as one architecture.
Full-service installation with no hidden costs
The platform sits between users and approved models, knowledge sources and tools. Access and identity are not treated as afterthoughts.
Employees, teams, external users or distributed offices.
NetBird, WireGuard, reverse proxy, segmentation and controlled exposure.
Keycloak or Authentik, SSO, MFA, groups and role-based permissions.
Shared interface, roles, model access, chats, knowledge and extensions.
Ollama, vLLM, LiteLLM or approved external model APIs.
RAG, vector database, object storage and connected document sources.
APIs, automations, business applications and custom-developed functions.
This is a modular reference design. Required components and their placement are derived from data classes, load, existing IT and operational requirements.
Open WebUI brings chat interface and RAG pipeline - but enterprise requirements demand more. We extend the frontend for your use cases.
We use the REST API for integrations and pipelines for custom model routing: The appropriate model is automatically selected based on input.
Input and output filters: Before sending to the model we can enrich prompts (RAG context), after the response we can perform compliance checks.
For UI customizations we develop custom Svelte components: Special input forms, visualizations, industry-specific chat layouts.
How we implement Open WebUI development in practice.
Different requests (code vs. text vs. analysis) need different models, but users shouldn't have to worry about it.
Pipeline that analyzes input and automatically routes to the optimal model (GPT-4, Claude, local Llama) - transparent to the user.
Internal documents (SharePoint, Confluence) should be searchable in chat without manually uploading them.
Automatic sync of enterprise documents into vector DB. With every question, relevant context from the knowledge base is injected.
LLM responses could contain sensitive info (PII, internal codes). Must filter before display.
Output pipeline that scans responses and masks or blocks sensitive data before displaying to user.
As an alternative to managed hosting in the data centre, WZ-IT provides the hardware, configures Open WebUI, and handles hardening, monitoring, updates, backup and technical support. Access can be limited to the internal network or enabled through VPN and existing identities.
€6,490 excl. VAT one-time · plus AI Cube Care at €349.90 excl. VAT / month

Specific answers on hosting, models, RAG, security, operations and licensing.
Yes. The agreed scope can cover the application, operating system and containers, monitoring, backups, updates and incident response. Human response times follow the selected service level; model usage and custom development are priced separately.
Yes. We install and operate Open WebUI in your data centre, on Proxmox, Kubernetes or in your cloud account. Before taking over operations, we assess access, backup, monitoring, technical debt and responsibility boundaries. Learn more about on-premises deployments.
Open WebUI supports Ollama and OpenAI-compatible APIs. This makes it possible to connect local inference systems such as Ollama or vLLM as well as approved external model providers. We select the gateway, models and routing according to data classes, quality, latency and cost.
No. When using external model APIs, the interface normally does not need its own GPU server. Large local language models are sized by model size, VRAM, concurrency and target latency. Where required, we design a managed GPU server.
We design ingestion, text extraction, embeddings, vector storage, metadata and source updates as one pipeline. For confidential content we also address permissions, citations, deletion concepts and the limits of subject-matter reliability.
Yes, provided a technical audit confirms that it can be operated reliably. We review the version, database, volumes, secrets, model connections, backups, network and recovery. Required remediation or migration is scoped transparently before ongoing operations begin.
Standard operations monitor availability and agreed technical states of the workload. Daily backups with seven-day retention are included in the standard offer. Additional databases, vector stores, GPU systems, longer retention and specific RTO/RPO targets are agreed to match the architecture.
For larger environments, we separate the application, database, storage and inference and design load balancing, multiple workloads, inference pools and recovery. High availability is not a generic hosting checkbox; it is designed around failure impact, RTO, RPO and load profile.
The current Open WebUI licence permits use with the original branding left intact. Removing or changing branding is subject to exceptions and may require an enterprise licence; the vendor's current licence is authoritative. We review technical implementation but do not provide legal advice.
Knowledge & guides
Read comparisons, installation guides and architecture articles independently of the specific WZ-IT service.
23.09.2026
Onyx is an open source platform that pulls company knowledge from many applications into a shared search index and generates AI answers with source citations...
14.09.2026
Public authorities manage large collections that have grown over years: decrees and circulars, procedural instructions, technical reports, expert opinions, statements, minutes. People looking for a...
31.08.2026
Article 50 of the AI Act has applied since 2 August 2026. Coverage of it reduces to a single sentence: companies must label their chatbots...
These solutions are often used together with Open WebUI
These solutions offer similar functionalities and can be evaluated together
These solutions are direct alternatives with similar use cases
No risk: worst case, you leave with a clearer understanding of your project than before.


“WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.”
“WZ-IT moved our studio infrastructure from decentralised individual devices to a central platform: every site is securely connected via VPN, new devices are onboarded automatically and an entire site is provisioned from a template, without manual steps on location. What impressed me most is the breadth and depth of their knowledge: Timo and Robin are not a typical IT provider who sets up a server and leaves. The two of them think their way into highly complex infrastructure and software topics, work through every requirement we put in front of them, and build networking, provisioning and operations so that everything fits together in the end. WZ-IT is an excellent partner for complex software, network and architecture projects.”

Steve Kirchner
Managing Director, nextGYM GmbH

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.