Cloud AI vs. self-hosted: which operating model?
Timo Wevelsiep•Updated: 04.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Compare cloud, dedicated, local and hybrid AI properly? WZ-IT evaluates data classes, quality, load and operating cost, then implements the selected model with gateway, monitoring and an exit path. Explore LLM hosting
AI operation is no longer a binary cloud-or-server choice. Public APIs, dedicated managed instances, self-hosting and hybrid combinations are all viable. This comparison covers data protection, quality, total cost, control and operational responsibility. As of August 2026.
Table of contents
- Two operating models for AI
- Data protection and control
- Cost: per token versus your own hardware
- Cloud AI and self-hosted compared
- When to use which model
- Dedicated and hybrid options
- Total cost instead of token price
- Decide with your own evaluation set
- How WZ-IT implements the platform
- Sources
Two operating models for AI
With cloud AI you use a language model as a service: your application sends requests to a provider's API, the model computes on their servers, the answer comes back. You need not provide hardware or operate anything - but model and data lie in someone else's hands.
With self-hosted AI, a selected model runs on controlled infrastructure (see What is local AI?). Your team or a managed-service provider may operate it. Actual control depends on administration, keys, data flows, exports and exit rather than the label alone.
Data protection and control
The central question is: which data is processed by whom, for what purpose and where? Beyond prompt and response, consider documents, embeddings, usage metadata, telemetry, error reports, support access and backups.
Self-hosting can avoid transfer to a model provider, but only if cloud features, web search, telemetry and external observability are controlled as well. Ollama, for example, documents a local-only setting that disables its cloud features. Test outbound connections rather than relying on product labels. International cloud transfers can be lawful under GDPR mechanisms; they still require a use-case-specific assessment.
Cost: per token versus your own hardware
Cloud AI is generally usage-based across input, output, embeddings and tools. It can be inexpensive at entry and for irregular demand. Variable costs grow with usage, while a self-hosted GPU mainly produces idle capacity when demand is low.
Self-hosted AI shifts cost into capacity and operations. Hardware, energy, redundancy, storage, backup, updates, monitoring and people all belong in the calculation. The useful metric is cost per successful task including human correction, not token or server cost alone.
Cloud AI and self-hosted compared
| Dimension | Public cloud API | Self-hosted AI |
|---|---|---|
| Data | processed under provider contract and controls | flows governed by the target architecture |
| Cost model | mainly variable with usage | capacity, platform and operations |
| Control | provider controls platform and portfolio | model, version and stack are selectable |
| Operations | low effort, integration remains | high effort, outsourceable as managed service |
| Scaling | strong for peaks | limited by provisioned capacity |
| Exit | depends on API, features and exports | easier with open formats, still needs testing |
When to use which model
The rule of thumb follows data and usage:
- Cloud AI where the data and contract fit, demand varies widely or a specific closed model is required.
- Self-hosted AI where data flows need tight control, base load is predictable or model and platform portability matter.
Dedicated and hybrid options
A dedicated managed instance combines stronger isolation and a defined region with outsourced operations. A hybrid platform exposes several model targets through a gateway. It can keep selected data classes local while using external models for specific capabilities or demand peaks.
LiteLLM can centralise deployments, fallbacks, timeouts and load balancing. It does not classify secrets automatically. The application must reliably provide tenant, data class and allowed targets, and fallback rules must not silently downgrade protection.
Total cost instead of token price
Include model and tool use, GPU idle time, storage, network, backup, high availability, replacement capacity, security updates, monitoring, human correction and exit. A mixed model is often economical: stable base load on controlled capacity, with selected external use for uncommon peaks or specialist tasks.
Decide with your own evaluation set
Define 30 to 100 representative tasks, expected results and prohibited data paths. Compare two to four variants on quality, grounded answers, latency, throughput, cost and operational risk. Include update, outage, load spike and rollback. This shows whether a smaller local model is sufficient or whether a dedicated or hybrid design adds measurable value.
How WZ-IT implements the platform
WZ-IT combines architecture and operations: LLM hosting for controlled endpoints, managed AI for the production stack and AI Cube for on-site scenarios. Gateway, identity, RAG, monitoring, backups, evaluation and recovery can be included.
The result is not a blanket cloud or local recommendation, but an architecture with explicit data paths, quality limits, costs and responsibilities. The open-source LLM stack explains its layers.
Sources
Rather have it operated?
You'd rather not run Local & Sovereign AI yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess an AI application and its infrastructure
We combine models, company knowledge, integrations, and operations into a reliable AI solution on your own or European infrastructure.
Frequently Asked Questions
Answers to the most important questions
With a public cloud API, the provider operates model and platform. With self-hosting, the organisation or its service provider operates a model in a controlled environment. Dedicated managed instances and hybrid architectures sit in between. They differ in data flows, model choice, scaling, cost, operations and exit capability.
Actual data flows matter more than labels. A cloud provider processes inputs and metadata according to its contract and technology. Self-hosting can avoid that transfer, but telemetry, model downloads, observability or support may still communicate externally. Both models need a reviewed data-flow diagram.
It depends on load, model, quality and operating effort. Cloud APIs often suit low or variable usage. Self-hosting can become attractive at predictable base load but must include hardware, energy, redundancy, updates, monitoring and people. Compare cost per successful task rather than token or GPU price alone.
Yes, operation is yours: hardware, model, stack, updates and monitoring. Cloud AI takes that off your hands, but you give up control and data. The effort can be outsourced as a managed service, so the advantages of self-operation stay without you having to carry the complexity yourself.
Yes. Via a gateway like LiteLLM, self-hosted and cloud models can be run behind a unified interface. This way you can process sensitive requests locally and uncritical ones at an external provider - the decision is made per use case, not across the board.
It depends on the task. Open models can be strong for extraction, classification, summarisation, RAG or code, while closed models may lead elsewhere. Use an internal evaluation set and assess quality thresholds, latency and cost per successful task.
More on Local & Sovereign AI
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for confidentiality professions
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council






