Architecture and migration
Sizing, target environment, data transfer, cutover and recovery are resolved before production operations.

WZ-IT plans, installs and operates Ollama as managed hosting, in your cloud or on premises. Depending on the target design, we also provide migration, secure network and identity integration, monitoring, backups, updates, integrations and further development.
Companies worldwide trust WZ-IT

Ollama is the inference layer for local language models. Instead of sending company data to external AI services, models such as Llama, Mistral, Gemma or Qwen run on your own infrastructure.
We operate Ollama for production use: with GPU tuning, model management, API gateways, monitoring, backup strategy and clean integration into Open WebUI, AnythingLLM, RAG pipelines and internal applications.
Ollama is easy to start, but production operations need more than a Docker container. We plan hardware, model sizes, VRAM budget, access paths, updates and observability around your use case.
The software is licensed under MIT. That makes Ollama a strong open foundation for sovereign AI infrastructure, provided security, operations and model governance are handled properly.
Models run on your own hardware, in your own data centre or on a server in a German data centre. With exclusively local models and services, prompts and outputs stay within controlled infrastructure.
We size VRAM, CPU, RAM and storage around model size, context window, user count and response-time requirements.
Clean handling of model versions, pull processes, approvals and rollbacks for reproducible AI environments.
On-premise, AI Cube, GPU server or private cloud: Ollama is operated so data protection, availability and access control fit together.
We connect Ollama to Open WebUI, AnythingLLM, LiteLLM, internal applications, automations and RAG pipelines.
Monitoring, updates, backups, model tests, security hardening and support turn a local demo into a reliable AI platform.
Ollama handles local model execution and forms the technical base for chat, RAG, agents and internal AI APIs.
We handle model changes, updates, GPU utilization, health checks and clear processes for production environments.
Access is provided through VPN, SSO, internal networks or API gateways. External models, integrations and update services are documented as separate data flows.
We do more than provide an application. WZ-IT designs the technical architecture, integrates network and identity, operates the agreed scope and develops integrations when the standard product is not enough.
Sizing, target environment, data transfer, cutover and recovery are resolved before production operations.
SSO, secure access, internal systems and existing security components are integrated appropriately.
Updates, backups, technical monitoring and response paths follow a transparent operational scope.
APIs, automation and custom extensions can be delivered beyond basic deployment.
The exact scope depends on the application, edition, infrastructure and criticality. Vendor licences and non-standard components are quoted separately.
AI applications need controlled model access, knowledge sources, permissions and observability in addition to the user interface. We design these data flows as one coherent stack.
Teams, business applications and API clients use defined interfaces and endpoints.
TLS, firewall rules, reverse proxies or private network paths are designed around the platform's exposure.
Local accounts, SSO, directories, service accounts and emergency access are connected through clear roles.
Application logic, model access, roles and approved capabilities run in a controlled environment.
Local or external models, GPU/CPU resources, routing, limits and cost controls.
Documents, databases, vector search and traceable data paths for RAG and search.
Reviewed tools, APIs, logging, metrics and technical quality controls.
Models, extensions and data sources are not enabled indiscriminately. Permissions, data exposure, cost and professional oversight are assessed per use case.
A clearly defined operating scope instead of an opaque hosting flat fee.
Compute, applications, storage and response are shown separately. You can see what ongoing operations include and which requirements need a technical assessment.
We also design custom Ollama architectures, integrations and migrations. Contact us for a technical assessment.
Select a technical starting point, additional storage and the required service level. We then validate the selection against the application, load profile, integrations and target architecture.
A workload is one compute instance with the applications agreed for it.
One standard app per workload is already included. Additional dedicated servers count as separate workloads.
Indicate additional storage for data, artifacts and backups.
We also implement custom backup schedules and retention periods. For example: daily backups retained for 30 days and one monthly backup retained for 12 months. Scope, storage requirements and additional costs are agreed separately.
Assessed individually; custom quote.
Standard hosting runs on a single instance without high availability. Availability for custom architectures is agreed separately.
Zero downtime on request
Distributed operation with several instances, a database cluster and rolling updates, so that a node failure causes no interruption. Only for applications suited to it; design and price after technical assessment.
This page covers installation and the agreed platform operations for Ollama. Custom code, the access layer and special confidentiality requirements remain clearly separated responsibilities that can be added when needed.
Maintain the application
Keep custom extensions, interfaces and dependencies controlled, updated and supportable.
from €699.90 excl. VAT / month
View serviceSecure access
Add identity-based access and private network paths as a separate operational layer.
from €349.90 excl. VAT / month
View serviceOperate confidentially
For professional secrecy holders, define contractual, access and infrastructure requirements in a dedicated operating model.
from €499.90 excl. VAT / month
View serviceThree relevant adjacent paths instead of a long list of further products.
All prices are net and exclude statutory VAT. The offers are addressed to businesses.
View the overall systemEnquiry
Briefly describe the current state and objective for Ollama. We assess infrastructure, integration, and ongoing operations.
The AI Cube covers the compact entry point. For larger models, rackmount or high concurrency, we design custom GPU systems.
Dedicated NVIDIA GPU servers in Germany including an agreed model, vLLM inference layer, Open WebUI and ongoing managed operations.
Compact NVIDIA GB10-based local AI server. Fully configured, hardened and prepared with a local model agreed in advance.
Answers to the most important questions
Ollama is MIT-licensed and free for local use. Costs only apply to the cloud models that Ollama additionally offers on its own infrastructure: usage-based after starter credits or on a monthly plan (as of October 2026). Self-hosting costs consist of servers or hardware and operations; we define the range based on model, number of users and response time.
When run locally, Ollama processes prompts and answers only on your own server; Ollama itself does not see this data. Cloud models are different: there, prompts go to the Ollama cloud. In production installations we disable the cloud features, i.e. cloud models and web search (OLLAMA_NO_CLOUD). We operate the server in your infrastructure or in a German data centre and sign a data processing agreement.
Ollama also runs on CPUs; for fluent answers across a team, a GPU with enough VRAM is the norm. VRAM determines which model size and context length fit. Parallel requests to the same model increase the required context memory, so we size by model, number of users and response time.
By default Ollama listens only locally on port 11434 and has no user management of its own. For team use we expose the service via OLLAMA_HOST and place a reverse proxy or gateway with authentication, TLS and access via VPN or internal network in front of it. Users usually work through Open WebUI.
Yes. Besides its own API, Ollama offers OpenAI-compatible endpoints for chat completions, completions, embeddings, models and the Responses API. Applications written for OpenAI can therefore connect to local models through a changed base URL.
Ollama fits single workstations, tests and small teams. When many users work at the same time, response times rise under load or a model has to be spread across several GPUs, vLLM is the better runtime. Since both offer OpenAI-compatible APIs, you can start with Ollama and migrate to vLLM later without rebuilding your applications.
29.09.2026
A service mailbox or ticket queue is usually sorted by hand: someone reads each request, assigns a topic, sets the priority and passes it to...
27.09.2026
Since 2025 a new group of open models has appeared that read document pages as images and output text, Markdown or JSON: olmOCR, PaddleOCR-VL, GLM-OCR,...
20.09.2026
XWiki ships without AI. Anyone who wants a chat with source references, writing assistance in the editor or semantic search across their own pages installs...
These solutions are often used together with Ollama
These solutions offer similar functionalities and can be evaluated together
These solutions are direct alternatives with similar use cases
“WZ-IT moved our studio infrastructure from decentralised individual devices to a central platform: every site is securely connected via VPN, new devices are onboarded automatically and an entire site is provisioned from a template, without manual steps on location. What impressed me most is the breadth and depth of their knowledge: Timo and Robin are not a typical IT provider who sets up a server and leaves. The two of them think their way into highly complex infrastructure and software topics, work through every requirement we put in front of them, and build networking, provisioning and operations so that everything fits together in the end. WZ-IT is an excellent partner for complex software, network and architecture projects.”

Steve Kirchner
Managing Director, nextGYM GmbH

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.