Sovereign AI means: models run on your own infrastructure, no data leaves the building. We build and operate local AI systems from one team - from GPU hardware through the open-source LLM stack to RAG knowledge systems. On these pages we place the tools honestly, show production operations and the path to GDPR-compliant AI.
What is local AI?
Local AI means: language models run on your own infrastructure instead of in the cloud. What that concretely means, why companies do it and when it makes sense.
Cloud AI vs. self-hosted
AI from the cloud or self-hosted? The comparison: data protection, cost per token versus your own hardware, control and effort - and when each model fits.
AI sovereignty for companies
AI sovereignty means control over data and models. Why 'a server in Frankfurt' isn't enough, what the legal framework demands - and the alternative.
Which LLM to self-host?
Which open language model to self-host? Model size by use case, the license question (Apache, MIT vs. Llama) and how to choose the right model.
Sizing GPU & VRAM
How much VRAM does an LLM need? The three components - model weights, KV cache and overhead - with rules of thumb, quantization and a worked example.
What is vLLM?
vLLM is a fast inference engine for LLMs: PagedAttention, continuous batching and an OpenAI-compatible API server. What it does and when it fits.
vLLM vs. Ollama
Ollama or vLLM for self-hosted LLMs? The comparison: simplicity vs. throughput, CPU vs. GPU, single user vs. production - and when each engine fits.
The open-source LLM stack
Self-hosting an LLM is not one tool but a stack: inference (vLLM), gateway (LiteLLM), observability (Langfuse), frontend and RAG - operated sovereignly.
What is LiteLLM?
LiteLLM is an OpenAI-compatible gateway across 100+ LLMs: one interface, central auth, budgets, rate limits and fallback. What it does and why.
What is Langfuse?
Langfuse is an open-source platform for LLM observability: tracing, evaluation and prompt management. What it does and why it matters for the EU AI Act.
What is RAG?
RAG connects an LLM with your own documents: how retrieval-augmented generation works, why it reduces hallucinations and where permissions become the crux.
Connect Open WebUI to Nextcloud (RAG with ACLs)
RAG on Nextcloud documents with real permissions: why ACLs are enforced before the vector search - identity, Nextcloud shares, Qdrant filter, POC to production.
Qdrant vs. pgvector
Qdrant or pgvector for RAG? The comparison: dedicated vector database versus Postgres extension, scaling, operations and filters - and when each fits.
Local AI for confidentiality professions
Lawyers, doctors, tax advisors: why cloud AI is not an option under professional secrecy (§203 StGB) and local AI is the compliance solution.
Processing documents with AI
Intelligent document processing with AI: how LLMs extend classic OCR, typical use cases and the sovereign, self-hosted path with Paperless-AI.
AI agents & automation
What sets AI agents apart from simple automation, how n8n orchestrates business processes and why sovereign, self-hosted operation is the safe path.
From one team
WZ-IT builds and runs Local & Sovereign AI in production for companies - design, build and operations from one team.
See LLM hosting →Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.
Timo Wevelsiep & Robin Zins
Managing Directors of WZ-IT
