What does a local AI server cost?
Timo Wevelsiep•Updated: 15.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Start with a transparent fixed price? The AI Cube Pro costs EUR 5,999 excluding VAT and includes hardware, local AI platform, an agreed model, hardening, initial setup, and five hours of support. Explore scope and configuration
The cost of a local AI server only becomes meaningful when hardware, platform, integration, and operations are separated. A retail computer price is not comparable to a production-ready multi-user platform. Conversely, a large integration project is unnecessary when a preconfigured standard already covers the use case.
Cost overview for getting started
| Item | Price or calculation | Scope |
|---|---|---|
| AI Cube Pro | EUR 5,999 excluding VAT, one-off | hardware, prepared AI platform, local model, hardening, initial setup, and five hours of support |
| Managed operations | from EUR 149.90 excluding VAT per month | ongoing updates, patches, and operations within the agreed service scope |
| Data sources and integrations | project-specific | for example Nextcloud, SharePoint, DMS, business software, RAG pipelines, or MCP servers |
| AI Cube Custom and clusters | project-specific | larger models, higher concurrency, special availability, or rack systems |
This provides a defined entry price. Total cost then depends on whether the built-in Open WebUI capabilities are sufficient or whether integrations and ongoing operations are required.
Four cost blocks
1. Hardware and base system
Available memory, memory bandwidth, and expected throughput determine the hardware class for local language models. Storage, network, and possibly redundancy add to it. A compact starting point may use one system; larger models or high concurrency require more systems or dedicated GPU servers.
2. AI platform setup
Production setup includes operating system, drivers, model runtime, user interface, model, identity, network, hardening, and functional testing. Organisations that build this internally should price the time and long-term knowledge required.
3. Workflow integration
Uploading documents manually to a knowledge interface is different from a continuously synchronised RAG pipeline. Nextcloud, SharePoint, DMS, or professional-software interfaces, source permissions, middleware, and automation depend on the project.
4. Ongoing operations
Updates, security review, monitoring, backup, recovery, model changes, and support remain throughout the lifecycle. They can be handled internally or through a managed service.
Concrete starting point: AI Cube Pro
The AI Cube Pro costs EUR 5,999 excluding VAT. It includes:
- hardware with 128 GB unified memory;
- Open WebUI and a local model runtime;
- a local model agreed, installed, and tested in advance;
- base setup, hardening, and functional testing;
- initial setup and five hours of support.
The system is usually ready within two weeks after configuration approval. Personal delivery and setup within Germany can be arranged.
What is scoped separately
- automated knowledge sources and custom RAG pipelines;
- PMS, DATEV, DMS, Nextcloud, or SharePoint connections;
- custom APIs, middleware, workflows, and MCP servers;
- VPN, NetBird, reverse proxies, or site connectivity;
- high availability, clusters, and larger GPU systems;
- ongoing operations and agreed service levels.
Optional managed operations start at EUR 149.90 excluding VAT per AI Cube and month. Scope and response model are agreed separately.
Compare three years, not the purchase price
Use equal assumptions for both paths:
Local platform: acquisition + setup + integration + 36 months of power and operations + replacement risk.
Cloud platform: 36 months of user licences or API consumption + integrations + administration + paid controls and regions.
Non-financial factors also matter: permitted data types, control over models and updates, offline operation, provider dependence, and internal operating effort.
When standard hardware is not enough
The AI Cube Pro is intended for organisations starting with local AI, testing their own workloads, and serving a defined user group. A larger deployment is appropriate when several large models run in parallel, contexts are very long, concurrency is high, or availability and redundancy need contractual commitments. WZ-IT then sizes AI Cube Custom or GPU servers from measured load rather than a model-name assumption.
The guide to sizing a local AI server by user load explains how account count, concurrency, and response length become a reproducible load profile.
Sources and further reading
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
The WZ-IT AI Cube Pro costs EUR 5,999 excluding VAT, including the preconfigured platform, an agreed local model, hardening, initial setup, and five hours of support. Custom integrations and ongoing operations are separate.
Include power, administration, updates, monitoring, backup, support, and potentially replacement hardware. Optional managed operations for the AI Cube start at EUR 149.90 excluding VAT per month.
That depends on user count, utilisation, required model quality, and operating scope. Cloud subscriptions are often cheaper for a few users; local systems can make more sense for stable use and strict control. Compare at least a three-year period.
Open WebUI lets users create personal and shared knowledge spaces. Automated data sources, a custom RAG pipeline, or interfaces to Nextcloud, SharePoint, DMS, and professional software are scoped separately.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for confidentiality professions
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





