15.08.2026
Local AI or ChatGPT Business? Cost, control, and use compared
ChatGPT Business and a local AI platform partly solve the same user problem: staff need a reliable AI workspace. Technically and organisationally, they are different...
A turnkey local AI platform for chat, document search and custom assistants. WZ-IT supplies the hardware, configures the platform and model, and handles monitoring, updates and support.
plus one-time provisioning and initial setup
Companies worldwide trust WZ-IT

AI Cube Pro · Personal delivery and on-site setup available in Germany
You receive a complete local AI platform that lets your business test its own use cases and run initial workloads in production. WZ-IT provides the hardware, platform, model and ongoing operations together.
AI Cube Pro Managed
Fully configured and managed local AI server for AI chat, document search, custom assistants and further applications inside the company network.
excl. VAT · plus one-time provisioning and initial setup · 6-month minimum term · then renewable month by month
Hardware, monitoring, updates and support are included in the monthly rate. Provisioning and initial setup are charged once. After the minimum term, you can continue, return or upgrade the AI Cube.
Before delivery, we clarify the use case, model, storage and access paths. We then configure and test the system in Germany.
Compact appliance
Unified memory for local models and long contexts
The exact SSD and model configuration is defined before provisioning based on data volume, model size and expected usage.
Open WebUI is the shared browser interface for employees. They can chat with local models, query their own documents and use assistants without switching between separate AI tools.
Preinstalled
Open WebUI combines chats, files, models and assistants. WZ-IT configures the interface on the Cube and hands it over ready for your team to use.
Chats, histories, files and model selection in an interface that teams can use without specialist tooling.
Authorised users can upload documents, create collections and use knowledge personally or together as a team.
Ollama or vLLM serves selected open-weight models directly on the AI Cube.
Configurable by WZ-IT
Identity, additional models and in-house tools are added only when your organisation needs them.
The AI Cube can prepare manuals, reports, policies and other approved information for internal AI search. Employees ask questions in natural language and receive answers with traceable source references.
Usable without an integration project
The knowledge feature is already configured. Authorised users can upload documents, create personal or shared knowledge areas and query them directly with source references.
Employees do not search for filenames but ask specific questions about their work.
Citations show which documents and passages an answer is based on.
Collections are structured around departments, tasks and approved user groups.
Extended by WZ-IT
When content should not be maintained manually, WZ-IT connects existing repositories and business systems, including synchronisation, permissions, OCR, metadata and quality testing.
RAG makes existing information easier to find but does not guarantee correct answers automatically. Document quality, text extraction, metadata and retrieval are therefore tested with representative questions.
This planning value describes responses being generated at the same time, not the number of user accounts. As usage grows, the infrastructure can be expanded with additional AI Cubes.
With the optimised standard configuration, one AI Cube Pro can generate up to ten responses at the same time. Additional employees can have their own accounts and use the system at different times.
The relevant number is how many people generate a response at the same time, not how many user accounts exist.
Long documents and extensive RAG contexts require more compute time than short chat questions.
For more parallel usage, we add further AI Cubes; larger models or specific requirements are covered by AI Cube Custom.
You receive a managed AI platform rather than an isolated appliance. WZ-IT prepares the platform and model, securely integrates the AI Cube with your environment and handles the agreed technical operations.
Configure, harden, test and document the operating system, drivers, Open WebUI, inference stack and agreed local model.
Networking, identity, knowledge sources and interfaces are connected to match the agreed scope.
Provide monitoring, updates and support while continuously adding new models, knowledge sources, interfaces and use cases.
Target architecture
The AI Cube can operate inside an isolated network or connect selectively to approved services. External access uses defined network and identity paths.
Local models and data sources without outbound model calls for particularly sensitive applications.
Employees and administrators connect through encrypted NetBird/VPN access without exposing the Cube publicly.
An upstream gateway enables external use with TLS, access control and clear network separation.
Which data stays local and which external services are reachable is defined and documented before commissioning.
AI Cube Pro Managed
at a monthly rate from EUR 899 excl. VAT · plus one-time provisioning and initial setup
Operating scope, maintenance windows, backup targets and responsibilities are documented before commissioning. Custom integrations, additional data sources and larger extensions remain separately plannable.
Technical monitoring of the agreed system components with documented alert routes.
Scheduled security and system updates for the operating system, containers and the agreed AI stack.
Backup management, checks and recovery procedures according to the agreed scope and storage target.
Essential is the base level. Business, Priority, Production and Critical add defined response times and a flexible operations allowance; Production and Critical cover round-the-clock P1 response.
We configure controlled data paths, permissions and administrative access. For managed operations, we agree the appropriate DPA and bind involved personnel to confidentiality in writing.
Enquiry
Tell us briefly how you want to use local AI and how many people will work with it. We will review the configuration, integration and suitable operating scope.
Defined first step
The RAG sprint adds a tested company knowledge set, citations, permissions path and subject-matter acceptance to the AI platform.
The AI Cube Pro covers the initial rollout and first production workloads. When models, concurrency, storage or availability requirements exceed that scope, we design the appropriate larger solution.
AI Cube Pro
For businesses introducing local AI for a team and starting with up to 10 parallel chats on one AI Cube Pro.

AI Cube Custom
For sustained higher concurrency, larger models, additional storage capacity, redundancy or specific network and operational requirements.

Connect multiple systems
As the number of simultaneously active users grows, additional AI Cubes can be added and requests distributed across them. For supported distributed workloads, two systems can also be linked directly through a 200 GbE QSFP connection.
Connecting systems does not automatically provide high availability or improve every workload. Software, model, parallelisation method and network topology must fit the workload; any HA commitment is designed separately.
NVIDIA GB10 class with 128 GB unified memory and a choice of 1 TB, 2 TB or 4 TB NVMe SSD. Compact appliance form factor, ConnectX-7 for linking several Cubes, 230 V operation. Fully configured and hardened with Linux, GPU drivers, Open WebUI, model runtime and the local model agreed in advance. Root access is yours.
Answers to the most important questions
Topics
An AI box combines local hardware with a usable AI platform. AI Cube Pro includes Open WebUI, the model runtime, an agreed local model, base hardening, functional testing, initial setup, and support. Businesses can use it to provide AI chat, document search, internal assistants, and other local AI workloads on their own network.
AI Cube Pro is based on validated ASUS/NVIDIA appliance hardware. For larger models, many concurrent users or special infrastructure requirements, we design AI Cube Custom with dedicated NVIDIA GPUs, multi-GPU, rackmount or custom networking.
Power consumption depends on the appliance or custom setup and the actual workload. AI Cube Pro is designed as a compact local AI box for office and enterprise environments; larger custom systems are assessed for power, cooling and site requirements in advance.
On AI Cube Pro the M.2 NVMe storage is replaceable. GPU and unified memory are part of the GB10 superchip and therefore fixed. As demand grows, we add further AI Cubes or move to AI Cube Custom with you. After the minimum term, the AI Cube can be continued, returned or replaced with a larger configuration.
It can, if only local models and local data sources are configured. Once external model providers, tools or integrations are enabled, data may leave the environment. We define and document these data paths during setup.
The AI Cube provides a technical foundation for local processing but does not replace a legal and organisational assessment. Model providers, data sources, permissions, logging, retention and administrative access must match the specific use case and be documented.
The AI Cube is technically suited to confidentiality-sensitive environments because models and data can be processed locally and access paths can be controlled. Whether a specific deployment meets legal and organisational requirements must be assessed for the individual use case.
Yes. Once the configuration is approved, we deliver the fully prepared AI Cube in under two weeks and commission it with you during an agreed remote appointment within that period. Hardware, Open WebUI, the model runtime and the agreed local model are configured, hardened and technically tested.
Yes. We handle the complete initial setup of the operating system, GPU drivers, container runtime, Open WebUI and model runtime, harden the base installation and verify the agreed local model before delivery.
The scope includes hardware, complete initial setup and hardening, the configured AI stack, a local model agreed and tested in advance, documentation, commissioning, onboarding and support. Populating knowledge bases, data migrations and custom integrations are quoted separately.
The AI Cube is prepared with Open WebUI as the interface and Ollama and/or vLLM for local inference. Open WebUI already includes knowledge collections and RAG. Local open-weight models, approved external models and additional vector databases are configured to match the use case.
Yes. Authorized users can upload documents in Open WebUI, create personal or shared knowledge bases and grant teams read or write access. Populating existing repositories, data migrations and automated integrations with DMS, Nextcloud or SharePoint are separate integration services.
Suitable sources include manuals, SOPs, reports, policies and project documentation from network drives, file shares, DMS platforms, Nextcloud, SharePoint, wikis or business systems. Automated synchronisation, text extraction, metadata and the mapping of existing permissions are designed for each source and implemented separately.
As a practical planning value, we size AI Cube Pro for up to 10 chats generating responses in parallel. Additional employees can have user accounts and use the system at different times. Actual speed depends on the configured model, context length, knowledge sources and response length.
Yes - depending on hardware configuration, multiple models can run in parallel. For intensive or parallel use, we recommend more powerful or customized hardware configurations.
Beyond chatbots and RAG systems: audio/video transcription, document indexing, data processing, code assistance, automation of internal processes - ideal for privacy-critical or compliance-relevant scenarios.
AI Cube Pro costs from EUR 899 excluding VAT per month plus one-time provisioning and initial setup. The minimum term is six months. Hardware, hardening, an agreed local model, commissioning, monitoring, updates and support are included. AI Cube Custom and custom integrations are quoted per project.
The minimum term is six months. After that, the AI Cube can be continued month by month, returned or replaced with a larger configuration. The hardware remains the property of WZ-IT throughout the term.
When data privacy, control, consistent performance, and long-term planning are important - e.g., with sensitive data, compliance requirements, or frequent AI use.
Yes. We support migration: data and model transfer, re-setup on your on-prem system - without external dependency.
The AI Cube is a managed local AI platform at your own site. A rented AI server runs in a data centre and is used as remote GPU infrastructure. The right option depends on data location, network connectivity and workload.
The operating system, GPU drivers, containers, Open WebUI, inference stack and models require regular updates. Agreed operations, monitoring, patch management and support are therefore already part of the monthly AI Cube service.
Yes. We connect the Cube to existing networks, identity services and data sources. For controlled external access, we can use NetBird or an upstream reverse proxy gateway without exposing the administration interface directly to the public internet.
Monitoring, patch management and technical support under the Essential service level are part of AI Cube Pro Managed. Business, Priority, Production and Critical add defined service windows, response paths and a flexible operations allowance. A separate replacement concept can also be agreed for time-sensitive systems.
WZ-IT handles fault analysis and the manufacturer's warranty process. Backup targets, recovery and any advance replacement are agreed to match the availability requirement; an immediate replacement device is not included without a separate service level.
We provide remote commissioning and support in German and English. On-site appointments or partner delivery can be planned for suitable projects.
AI Cube Pro is based on validated ASUS/NVIDIA appliance hardware. Configuration, installation of the local AI stack, testing and commissioning are done by us in Germany. For AI Cube Custom we additionally select components to match the workload and integrate them at our site.
Yes - we offer a reseller program with attractive purchasing conditions, technical support, and optional white-label license. Ideal for system integrators and IT service providers.
15.08.2026
ChatGPT Business and a local AI platform partly solve the same user problem: staff need a reliable AI workspace. Technically and organisationally, they are different...
25.06.2026
How local RAG systems unlock case knowledge, cite sources and preserve matter-level permissions in every answer.
10.05.2026
Three unauthenticated API calls. No login, no exploit framework, no privilege escalation. Three POST requests to a default port, and the machine's memory is on...
28.04.2026
DGX Spark and the WZ-IT AI Cube are not fundamentally different performance classes today. Both use NVIDIA's GB10 class with 128 GB unified memory. Comparing...
24.11.2025
OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a...
09.11.2025
Local AI inference means running a language or multimodal model on owned hardware rather than through a public API. Organisations gain control over data paths,...
08.11.2025
The use of Large Language Models (LLMs) such as GPT-4, Claude or Llama has evolved from experimental applications to mission-critical tools in recent years. However,...
No risk: worst case, you leave with a clearer understanding of your project than before.


“WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.”
From local AI integration to architecture, data sovereignty and ongoing operations.
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.
