Local AI does not begin with an abstract architecture. The AI Cube Pro is our preconfigured starting point with Open WebUI, a local model, hardening, and initial setup. These guides cover data protection, models, cost, sizing, RAG, and secure operations and show when standard hardware fits or a larger deployment is appropriate.
What is local AI?
Local AI explained: trust boundaries, architecture, hardware, data protection and the path from pilot to production on-premises or self-hosted AI.
Cloud AI vs. self-hosted
Cloud AI, dedicated operation, self-hosting or hybrid? Compare data protection, control, quality, TCO and operations with a practical decision matrix.
Private ChatGPT for business
Private ChatGPT for business: compare a cloud workspace, EU platform, private cloud and local AI by data flows, knowledge, cost and operations.
AI sovereignty for companies
AI sovereignty means control over data, models and operations. A practical framework for assessing cloud, managed AI and on-premises AI.
Chatbot or knowledge navigator?
Support chatbot, guided knowledge search, or simply a good full-text search? Four options compared - including the cases where none of them pays off.
Which LLM to self-host?
Which LLM should you self-host? Choose by task, quality, licence, language, VRAM, throughput and operations using a practical evaluation matrix.
Sizing GPU & VRAM
How much VRAM does an LLM need? The three components - model weights, KV cache and overhead - with rules of thumb, quantization and a worked example.
Inference vs. Training
Training creates models, inference uses them - and drives your costs. Why most companies need no training, which GPU is enough and when RAG suffices.
What does a local AI server cost?
Calculate local AI server cost across hardware, setup, users, operations, integration, and a three-year comparison with cloud subscriptions.
Buy or rent an AI server?
Buy or rent an AI server: compare cost, term, hardware replacement, operations, upgrades, and TCO for appliances, managed servers, and own infrastructure.
Size a local AI server by users
Size a local AI server by concurrency, response length, context, model, and measured load as a reliable basis for business hardware selection.
LLM models on 128 GB unified memory
LLMs for 128 GB unified memory: supported GB10 models, quantisation, KV cache, context, and AI Cube selection with practical model examples.
Connect AI Cubes with ConnectX-7
Connect AI Cubes through ConnectX-7: direct topologies, NVIDIA Sync, load balancing, distributed inference, and technical GB10 cluster limits.
Configure ASUS Ascent GX10 for business
Configure ASUS Ascent GX10 for business with local models, Open WebUI, hardening, users, network, backup, and production operations.
Configure NVIDIA DGX Spark for business
Configure NVIDIA DGX Spark as a local AI platform with models, Open WebUI, users, network hardening, backup, monitoring, and operations.
Configure Acer Veriton GN100 for business
Configure Acer Veriton GN100 for local AI with DGX OS, models, Open WebUI, users, network hardening, backup, monitoring, and operations.
Configure Dell Pro Max with GB10 for business
Configure Dell Pro Max with GB10 for production: appliance mode, models, Open WebUI, users, hardening, backup, monitoring, and operations.
Configure Gigabyte AI TOP ATOM for business
Configure Gigabyte AI TOP ATOM as a local AI platform with models, Open WebUI, ARM64, users, hardening, backup, ConnectX-7, and monitoring.
Configure HP ZGX Nano G1n for business
Configure HP ZGX Nano G1n for production with a local model, Open WebUI, ZGX Toolkit, users, hardening, backup, monitoring, and operations.
Configure Lenovo ThinkStation PGX for business
Configure Lenovo ThinkStation PGX for local AI with DGX OS, models, Open WebUI, users, network hardening, backup, monitoring, and operations.
Configure MSI EdgeXpert for business
Configure MSI EdgeXpert for production with local models, Open WebUI, NVIDIA AI Enterprise, users, hardening, backup, monitoring, and operations.
What is vLLM?
vLLM is a fast inference engine for LLMs: PagedAttention, continuous batching and an OpenAI-compatible API server. What it does and when it fits.
vLLM vs. Ollama
Ollama or vLLM for self-hosted LLMs? The comparison: simplicity vs. throughput, CPU vs. GPU, single user vs. production - and when each engine fits.
The open-source LLM stack
Self-hosting an LLM is not one tool but a stack: inference (vLLM), gateway (LiteLLM), observability (Langfuse), frontend and RAG - operated sovereignly.
What is LiteLLM?
LiteLLM is an OpenAI-compatible gateway across 100+ LLMs: one interface, central auth, budgets, rate limits and fallback. What it does and why.
What is Langfuse?
Langfuse is an open-source platform for LLM observability: tracing, evaluation and prompt management. What it does and why it matters for the EU AI Act.
Run Open WebUI as a production appliance
Run Open WebUI as a production local AI platform with identity, knowledge, models, persistent data, backup, updates, monitoring, and secure access.
What is RAG?
RAG connects an LLM with your own documents: how retrieval-augmented generation works, why it reduces hallucinations and where permissions become the crux.
Knowledge transfer during employee transitions
Plan knowledge transfer when employees leave: select topics, conduct interviews, review the content and make approved expertise accessible to successors.
Connect Open WebUI to Nextcloud (RAG with ACLs)
RAG on Nextcloud documents with real permissions: why ACLs are enforced before the vector search - identity, Nextcloud shares, Qdrant filter, POC to production.
Qdrant vs. pgvector
Qdrant or pgvector for RAG? The comparison: dedicated vector database versus Postgres extension, scaling, operations and filters - and when each fits.
RAG with permissions
Access rights in RAG systems: why the prompt is not a security boundary, where the filter belongs, and what a post-filter leaks through its hit count.
RAG with Nextcloud, SharePoint, and DMS
Connect Nextcloud, SharePoint, or DMS to local AI with controlled synchronisation, permissions, deletion, and a production RAG architecture.
The EU AI Act for companies
EU AI Act for companies: roles, risk levels, deadlines after the 2026 AI Omnibus and a practical implementation plan for governance and operations.
AI agents: permissions and approvals
Whose account does an AI agent act under? Service account versus delegated identity, a tool catalogue instead of full access, approvals and logging.
AI assistants and the works council
Internal AI assistants and German works councils: section 87 co-determination, data protection, logging and technical rules for a works agreement.
GDPR-compliant AI: assessment criteria
Assess GDPR-compliant AI in an organisation: purpose, legal basis, data paths, deletion, permissions, providers, and technical controls in practice.
Provide secure remote access to local AI
Provide local AI securely to remote staff and sites: compare VPN, NetBird, and reverse-proxy access without exposing model or database ports.
Local AI for professional secrecy holders
AI under section 203 StGB: requirements for lawyers, doctors and tax advisers, operating models, technical controls and a practical assessment path.
Processing documents with AI
Intelligent document processing with AI: how LLMs extend classic OCR, typical use cases and the sovereign, self-hosted path with Paperless-AI.
AI agents & automation
What sets AI agents apart from simple automation, how n8n orchestrates business processes and why sovereign, self-hosted operation is the safe path.
From one team
WZ-IT builds and runs Local AI for Business in production for companies - design, build and operations from one team.
Explore the AI Cube Pro →Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.