Architecture and migration
Sizing, target environment, data transfer, cutover and recovery are resolved before production operations.
WZ-IT plans, installs and operates PrivateGPT as managed hosting, in your cloud or on premises. Depending on the target design, we also provide migration, secure network and identity integration, monitoring, backups, updates, integrations and further development.
The following are trademarks of their respective owners: PrivateGPT (Zylon by PrivateGPT). WZ-IT is an independent service provider and has no business, partnership, or contractual relationship with these companies. We offer independent migration, installation, hosting, and operations services.
PrivateGPT provides building blocks for document-based RAG applications using large language models. Document processing, retrieval and model access can run locally; when remote LLM or embedding providers are used, the content sent to them is processed externally.
The platform supports local and remote model providers. We therefore design data flows, model access, vector storage and permissions for the required protection level instead of treating self-hosting as automatic privacy compliance.
We install, host and operate PrivateGPT either on-premise or on agreed infrastructure in Germany. Local and external model and embedding services are designed as traceable data flows.
We provide 24/7 monitoring, backups and maintenance for your PrivateGPT instance. Human response times and support coverage follow the selected service level. Including support for integrating your preferred LLM providers, document ingestion, and RAG optimization.
Documents, embeddings and model requests can remain in your environment with an entirely local configuration. External providers are assessed and approved as separate data flows.
RAG pipeline with document parsing, splitting, metadata, embeddings and chunk retrieval. Result quality and source grounding are tested against the concrete corpus.
FastAPI-based architecture with OpenAI-compatible endpoints. Required parameters, streaming, tool and client capabilities are tested by version and adapted where needed.
Can be operated completely offline and locally - on-premise, in your private cloud, or air-gapped environment. No dependency on external cloud services.
Supports local providers (Ollama, LlamaCPP) and remote providers (OpenAI, Azure OpenAI, Sagemaker, etc.). Flexibly switch between different models.
Document processing with parsing, chunking, metadata and embeddings. Result quality and supported formats depend on version, parsers and configuration.
A local design can control storage locations, identities and network paths. Legal compliance additionally depends on the use case and connected services.
The publicly available source code is licensed under Apache 2.0 and can be reviewed and adapted. Models and other dependencies are assessed separately.
Based on LlamaIndex framework. Free choice between different LLM providers, vector databases, and deployment models without dependency on individual vendors.
Professional installation on your infrastructure - on-premise, cloud or hybrid
In your data center
AWS, Azure, Hetzner & more
Advanced architecture after technical and licence assessment
Search confidential documents in a controlled environment; data flows, model access and telemetry are configured for your architecture
Controllable data flows for organisations with elevated protection needs; legal and domain requirements are assessed per project
Production-ready RAG with intelligent document parsing, embeddings generation and context-based chunk retrieval for precise answers
Flexible use of local LLMs (Ollama, LlamaCPP) or remote providers (OpenAI, Azure OpenAI, Gemini) depending on requirements
Fully operable offline for layered security controls requirements, isolated networks and air-gapped environments without external dependencies
FastAPI-based architecture with OpenAI-compatible API for seamless integration into existing tools and workflows without code changes
Secure access and access control for your installation
WireGuard, NetBird or Tailscale
Directly or through an upstream identity layer
Depends on application, edition and identity provider
Fail2Ban, Rate Limiting, IP Whitelisting
We set up secure VPN access to your installation - ideal for remote work and external employees.
Full-service installation with no hidden costs
PrivateGPT enables local LLM usage with your own documents. We integrate it into your existing infrastructure and extend capabilities.
Via the API we connect PrivateGPT to your systems: Chat completion, document ingest, and context retrieval can be controlled programmatically.
We develop custom ingestion pipelines: Special document types (CAD, contracts, code) are processed with optimized chunking strategies.
The LlamaIndex backend allows deep customization: Custom retrievers, rerankers, and embedding models for domain-specific search.
How we implement PrivateGPT development in practice.
Highly sensitive environments (government, defense) cannot have cloud connections but need LLM capabilities.
Fully offline-capable setup with local models, own embeddings, and airgapped update mechanisms.
Lawyers need to search thousands of documents for relevant passages. Manual review takes weeks.
PrivateGPT with legally optimized chunking and retrieval. Questions in natural language, answers with source references.
Executives receive hundreds of emails daily. Summaries and prioritization are needed.
PrivateGPT integration that analyzes, summarizes, and categorizes emails by urgency - fully on-premise.
We do more than provide an application. WZ-IT designs the technical architecture, integrates network and identity, operates the agreed scope and develops integrations when the standard product is not enough.
Sizing, target environment, data transfer, cutover and recovery are resolved before production operations.
SSO, secure access, internal systems and existing security components are integrated appropriately.
Updates, backups, technical monitoring and response paths follow a transparent operational scope.
APIs, automation and custom extensions can be delivered beyond basic deployment.
The exact scope depends on the application, edition, infrastructure and criticality. Vendor licences and non-standard components are quoted separately.
AI applications need controlled model access, knowledge sources, permissions and observability in addition to the user interface. We design these data flows as one coherent stack.
Teams, business applications and API clients use defined interfaces and endpoints.
TLS, firewall rules, reverse proxies or private network paths are designed around the platform's exposure.
Local accounts, SSO, directories, service accounts and emergency access are connected through clear roles.
Application logic, model access, roles and approved capabilities run in a controlled environment.
Local or external models, GPU/CPU resources, routing, limits and cost controls.
Documents, databases, vector search and traceable data paths for RAG and search.
Reviewed tools, APIs, logging, metrics and technical quality controls.
Models, extensions and data sources are not enabled indiscriminately. Permissions, data exposure, cost and professional oversight are assessed per use case.
A clearly defined operating scope instead of an opaque hosting flat fee.
Compute, applications, storage and response are shown separately. You can see what ongoing operations include and which requirements need a technical assessment.
We also design custom PrivateGPT architectures, integrations and migrations. Contact us for a technical assessment.
One managed standard PrivateGPT application is included in the Starter workload. Every service level also includes flexible expert time for planned work during regular service hours. Select compute, additional applications, storage and the appropriate service level.
A workload is one compute instance with the applications agreed for it.
One standard app per workload is already included. Additional dedicated servers count as separate workloads.
€79.90 per started TB and month, including daily encrypted offsite backup with 7-day retention.
From fully managed GPU servers to compact AI Cubes - we provide the ideal infrastructure for your local LLM applications.
Powerful GPU servers with dedicated hardware for compute-intensive LLM workloads. Fully managed, scalable, and optimized for maximum performance.
Compact AI workstation for local LLM inference. Perfect for office environments, with top-tier performance and absolute data sovereignty.
Whether you run PrivateGPT in-house or need to host confidentiality-professional data §203-ready - we build, operate and maintain PrivateGPT on an encrypted on-site server. Data never leaves the building in cleartext.
See PrivateGPT on-premise
These solutions are often used together with PrivateGPT
These solutions offer similar functionalities and can be evaluated together
These solutions are direct alternatives with similar use cases
No risk: worst case, you leave with a clearer understanding of your project than before.


“WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.”
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.
Timo Wevelsiep & Robin Zins
Managing Directors of WZ-IT
