Sovereign AI means: models run on your own infrastructure, no data leaves the building. We build and operate local AI systems from one team - from GPU hardware through the open-source LLM stack to RAG knowledge systems. On these pages we place the tools honestly, show production operations and the path to privacy-focused AI.
What is local AI?
Local AI explained: trust boundaries, architecture, hardware, data protection and the path from pilot to production on-premises or self-hosted AI.
Cloud AI vs. self-hosted
Cloud AI, dedicated operation, self-hosting or hybrid? Compare data protection, control, quality, TCO and operations with a practical decision matrix.
AI sovereignty for companies
AI sovereignty means control over data, models and operations. A practical framework for assessing cloud, managed AI and on-premises AI.
Chatbot or knowledge navigator?
Support chatbot, guided knowledge search, or simply a good full-text search? Four options compared - including the cases where none of them pays off.
Which LLM to self-host?
Which LLM should you self-host? Choose by task, quality, licence, language, VRAM, throughput and operations using a practical evaluation matrix.
Sizing GPU & VRAM
How much VRAM does an LLM need? The three components - model weights, KV cache and overhead - with rules of thumb, quantization and a worked example.
Inference vs. Training
Training creates models, inference uses them - and drives your costs. Why most companies need no training, which GPU is enough and when RAG suffices.
What is vLLM?
vLLM is a fast inference engine for LLMs: PagedAttention, continuous batching and an OpenAI-compatible API server. What it does and when it fits.
vLLM vs. Ollama
Ollama or vLLM for self-hosted LLMs? The comparison: simplicity vs. throughput, CPU vs. GPU, single user vs. production - and when each engine fits.
The open-source LLM stack
Self-hosting an LLM is not one tool but a stack: inference (vLLM), gateway (LiteLLM), observability (Langfuse), frontend and RAG - operated sovereignly.
What is LiteLLM?
LiteLLM is an OpenAI-compatible gateway across 100+ LLMs: one interface, central auth, budgets, rate limits and fallback. What it does and why.
What is Langfuse?
Langfuse is an open-source platform for LLM observability: tracing, evaluation and prompt management. What it does and why it matters for the EU AI Act.
What is RAG?
RAG connects an LLM with your own documents: how retrieval-augmented generation works, why it reduces hallucinations and where permissions become the crux.
Connect Open WebUI to Nextcloud (RAG with ACLs)
RAG on Nextcloud documents with real permissions: why ACLs are enforced before the vector search - identity, Nextcloud shares, Qdrant filter, POC to production.
Qdrant vs. pgvector
Qdrant or pgvector for RAG? The comparison: dedicated vector database versus Postgres extension, scaling, operations and filters - and when each fits.
RAG with permissions
Access rights in RAG systems: why the prompt is not a security boundary, where the filter belongs, and what a post-filter leaks through its hit count.
The EU AI Act for companies
EU AI Act for companies: roles, risk levels, deadlines after the 2026 AI Omnibus and a practical implementation plan for governance and operations.
AI agents: permissions and approvals
Whose account does an AI agent act under? Service account versus delegated identity, a tool catalogue instead of full access, approvals and logging.
AI assistants and the works council
Internal AI assistants and German works councils: section 87 co-determination, data protection, logging and technical rules for a works agreement.
Local AI for confidentiality professions
AI under section 203 StGB: requirements for lawyers, doctors and tax advisers, operating models, technical controls and a practical assessment path.
Processing documents with AI
Intelligent document processing with AI: how LLMs extend classic OCR, typical use cases and the sovereign, self-hosted path with Paperless-AI.
AI agents & automation
What sets AI agents apart from simple automation, how n8n orchestrates business processes and why sovereign, self-hosted operation is the safe path.
From one team
WZ-IT builds and runs Local & Sovereign AI in production for companies - design, build and operations from one team.
See LLM hosting →Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.
Timo Wevelsiep & Robin Zins
Managing Directors of WZ-IT
