What is local AI? Models on your own infrastructure
Timo Wevelsiep•Updated: 18.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Start using local AI in under two weeks? The AI Cube Pro arrives fully preconfigured with Open WebUI, an agreed local model, hardening, initial setup, and support. Explore the AI Cube Pro
Local AI is more than a language model running on a GPU. It becomes production-ready only with a defined trust boundary, identity, knowledge sources, observability, updates and recovery. This article explains the architecture and when on-premises, dedicated self-hosting or hybrid operation makes sense. As of August 2026.
Table of contents
- Cloud AI or local AI
- What "local" actually means
- Why companies run it locally
- What belongs to local AI
- When local AI makes sense
- Where local systems still communicate externally
- The preconfigured starting point with the AI Cube Pro
- From pilot to production
- How WZ-IT implements local AI
- Sources
Cloud AI or local AI
Most people know AI through cloud services: an application sends a request to an externally operated model endpoint. Depending on provider, product and configuration, content and metadata are processed in the agreed region. This is convenient and scales quickly, but delegates part of the technical control.
Local AI moves inference into a controlled environment. Users can keep the same chat or API experience, while the organisation decides model, version, network paths, access and retention. This creates control but also operational responsibility.
What "local" actually means
"Local" does not necessarily mean "in your own server room". What is meant is: on infrastructure you control. That can be:
- a GPU server in your own data center or server room (on-premise),
- a virtual machine on Proxmox or bare metal,
- a server at a European hoster that is under your control.
The decisive concept is the defined trust and operating boundary. On-premises normally means hardware at the organisation's site. Self-hosted means that the stack is operated by or for the organisation. Dedicated infrastructure can sit at a European hosting provider. These terms should not be used interchangeably because access, responsibility and resilience differ.
Why companies run it locally
Three reasons drive the switch:
- Data protection and sovereignty - data paths, access and retention can be constrained more tightly. Whether external transfers disappear entirely depends on the whole platform. This strengthens the technical basis for data protection and AI sovereignty but does not replace legal assessment.
- Cost control - capacity cost is predictable, while hardware, energy, redundancy and operations must also be included. Local inference can suit stable base load; cloud can remain cheaper for low or highly variable demand.
- Independence - no lock-in to the price, model or license changes of a single provider. You decide which model runs when.
The trade-off in detail - when cloud, when your own hardware - is shown in Cloud AI vs. self-hosted.
What belongs to local AI
Local AI is a platform with several layers:
- Compute and storage - GPU, CPU, VRAM, model storage, document storage and backup.
- Inference server - for example Ollama for compact setups or vLLM for throughput-oriented APIs.
- Gateway - common endpoints, model routing, limits and fallbacks.
- Identity and network - SSO, roles, secrets, segmentation and controlled administration.
- Application and knowledge - UI, APIs, RAG, citations and permission checks before retrieval.
- Operations - metrics, data-minimised traces, evaluation, updates, rollback, backup and incidents.
Only this interplay turns a model demo into a multi-user production service with defined rights and availability.
When local AI makes sense
Local AI pays off especially when at least one of these applies: you process sensitive or regulated data (law, health, public sector, industry), you have continuous, high usage where token costs weigh in, or you want to be independent of a single US provider.
For sporadic and suitable use, a cloud service may remain the simpler entry. Local AI becomes particularly relevant when data paths need tight boundaries, systems must integrate with existing identity and networks, or base load is predictable. Decide through a pilot rather than a blanket privacy or cost claim.
Where local systems still communicate externally
“The model is local” does not mean the platform is offline. Typical outbound paths include model and container downloads, telemetry, crash reporting, web search, external embeddings, cloud observability, email, OCR, speech services, remote support and off-site backups. These may be useful and lawful, but should be visible, approved and constrained.
Ollama, for example, provides a setting to disable its cloud features. Egress policies and outbound-connection tests provide additional assurance.
From pilot to production
- Define use case, user groups, data classes and measurable quality targets.
- Compare two or three exact model releases on real tasks and documents.
- Measure VRAM, context, parallelism, latency and peak demand.
- Include identity, permissions, RAG and logging in the pilot.
- Test updates, failure, backup restore and model rollback.
- Only then size hardware and select a production service level.
This avoids buying oversized hardware for an unsuitable model or building a strong demo without an operating path.
The preconfigured starting point with the AI Cube Pro
Organisations without an existing AI platform do not need to assemble every layer separately. The AI Cube Pro combines hardware with 128 GB of unified memory, Open WebUI, a local model runtime, and a model agreed in advance. WZ-IT handles base setup, hardening, functional testing, initial setup and support. The system is usually ready within two weeks after configuration approval.
Staff can then use a familiar chat interface with the local model and create personal or shared knowledge spaces in Open WebUI. Automated data sources, a custom RAG pipeline, professional-software interfaces, VPN or NetBird access, and process automation are not blanket standard features. They are scoped as separate integrations when required.
AI Cube Pro Managed costs from EUR 899 excluding VAT per month, plus one-time provisioning and initial setup, with a six-month minimum term. Hardware, monitoring, updates and support are included. If model size, user count, or availability requirements exceed this platform, the same decision path leads to additional AI Cubes, AI Cube Custom, or a dedicated GPU server.
How WZ-IT implements local AI
WZ-IT can integrate the platform into existing infrastructure or operate it as managed AI. The AI Cube supports compact on-premises scenarios; GPU servers and LLM hosting cover larger or centralised model services.
The work goes beyond starting a model: network, identity, model server, knowledge sources, monitoring, backup and further development become one operable system. An internal AI assistant with sources and permissions can build on top.
If the main objective is a centrally managed employee chat, private ChatGPT for business compares a cloud workspace, European platform, dedicated self-hosting and local operation.
Sources
- Ollama documentation: local operation, cloud controls and parallelism
- vLLM documentation: OpenAI-compatible inference server
- LiteLLM documentation: routing and load balancing
- European Commission: data-protection obligations and DPIAs
- German Data Protection Conference: guidance on AI and data protection
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
Local AI means that inference and, where applicable, knowledge retrieval run inside a controlled environment such as on-premises, an organisation's data centre or dedicated infrastructure. Whether data actually stays within that boundary also depends on telemetry, web search, model downloads, observability, backups and support access.
With cloud services, input is processed through the provider's infrastructure; data regions, retention, and administrative controls depend on the selected plan. With local AI, the model runs inside your defined environment. External paths for updates, web search, remote support, or optional cloud models still need a separate assessment.
For production operation of larger language models you usually need a GPU with sufficient VRAM. Small models also run on CPU or modest hardware. The right sizing depends on model size, desired throughput and number of users - from a single GPU server to a small cluster.
Local AI can reduce external transfers and improve technical control, but it is not automatically GDPR-compliant. Legal basis, purpose limitation, minimisation, access, deletion, processing agreements and potentially a data protection impact assessment still need to be considered. The complete processing operation matters, not just the model server.
That depends on usage, model, availability, and operating effort. Local AI has acquisition and operating costs, while cloud offerings are generally billed by user or consumption. A useful comparison needs the same user count, quality target, and a period of at least three years.
Closed cloud models cannot simply be installed locally. Many models are available under open or community licences. Quality and usage rights vary by exact release, so review the model card and licence and test realistic tasks before production use.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Knowledge transfer during employee transitions
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- Private ChatGPT for business
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for professional secrecy holders
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Buy or rent an AI server?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





