Local AI Inference with the AI Cube: Run AI Inside Your Own Network

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

View the WZ-IT AI Cube - local AI server with 128 GB unified memory, complete initial setup, hardening, Open WebUI and an agreed model. Schedule an initial consultation
Local AI inference means running a language or multimodal model on owned hardware rather than through a public API. Organisations gain control over data paths, model selection and cost structure. They also take responsibility for hardware, software, permissions, updates and operations.
The WZ-IT AI Cube combines these components in a compact local AI server. It is not delivered as an empty developer machine, but fully configured, hardened and prepared with a local model agreed in advance.
Table of contents
- What is processed locally
- AI Cube hardware
- What is configured before delivery
- Open WebUI as the workspace
- RAG and internal documents
- Sizing and limits
- Self-operation or managed service
- Cost and decision
What is processed locally
In a fully local configuration, model, chat workspace, knowledge index and database run inside the organisation's own network. Prompts, uploaded files and model responses do not need to be transferred to an external model provider.
Whether there are truly no external connections depends on configuration. Cloud models, web search, external tools or update services create additional data paths. We define and document these connections during setup.
Local processing is particularly relevant when:
- confidential documents are processed,
- professional confidentiality or customer requirements apply,
- public model APIs are not approved,
- fixed local capacity is preferred over variable API usage,
- custom integrations should avoid provider lock-in.
AI Cube hardware
The standard configuration uses a compact ASUS/NVIDIA appliance in the GB10 class:
| Feature | Standard configuration |
|---|---|
| Compute platform | NVIDIA GB10 Grace Blackwell |
| Processor | 20-core Arm |
| Memory | 128 GB LPDDR5x unified memory |
| Storage | 1, 2 or 4 TB NVMe depending on configuration |
| Networking | 10 GbE and ConnectX-7 |
| Form factor | Compact desktop appliance |
Unified memory is shared between CPU and GPU. It allows quantised models to fit into a memory class unavailable on many individual graphics cards. Actual speed still depends on model architecture, quantisation, context and concurrency.
What is configured before delivery
Before ordering, we discuss the workload and agree the local model. A larger model is not automatically better: for some tasks, a smaller model provides lower latency and more concurrency.
The AI Cube is then prepared for plug-and-play use:
- Linux, GPU drivers and container runtime fully configured,
- base installation hardened and prepared for operation,
- Open WebUI installed and configured,
- Ollama and/or vLLM set up for the workload,
- agreed local model installed and verified,
- delivered configuration documented.
Five hours for integration into your environment are also included. This time can cover network, identity, an initial knowledge source or model parameters. Larger integrations are planned separately.
Open WebUI as the workspace
Open WebUI provides a browser-based workspace. Employees receive a familiar chat interface while administrators control models, users, knowledge collections and connections.
Depending on configuration, it can provide:
- chat and history with local models,
- multiple approved models in one interface,
- file uploads and knowledge collections,
- roles, groups and OIDC integration,
- custom assistants with system prompts and knowledge,
- OpenAI-compatible interfaces for internal applications.
External cloud models can optionally be connected alongside local models. Sensitive workloads can also run with local models only.
RAG and internal documents
Retrieval-augmented generation adds relevant content from internal documents to a model. The process involves several stages:
- Documents are extracted and split into useful sections.
- Embeddings represent these sections for search.
- Relevant passages are retrieved for a question.
- The model generates an answer from that context.
- Citations make the answer traceable.
Open WebUI includes knowledge features for initial use cases. For large, frequently changing or permission-sensitive collections, we build separate pipelines, vector databases, middleware and interfaces to Nextcloud, BookStack, document management or line-of-business applications.
Sizing and limits
Reliable sizing considers at least:
- model and quantisation,
- prompt, context and answer length,
- concurrent active users,
- interactive or batch use,
- RAG, image processing or tool calls,
- availability and recovery requirements.
One AI Cube often fits local assistants, document search and inference for a limited team. For high concurrency, HA, rackmount or dedicated GPUs, we design custom GPU servers.
Two AI Cubes can be linked through ConnectX-7 and supported workloads can be distributed. Linking systems does not automatically create high availability or accelerate every model.
Self-operation or managed service
You receive root access and can operate the AI Cube yourself. Operations include:
- security and operating system updates,
- updates for Open WebUI and the model runtime,
- model changes and compatibility checks,
- monitoring of hardware, services and storage,
- backups of configuration, database and knowledge collections,
- controlled recovery.
WZ-IT can optionally take over these tasks. Scope and response times are agreed for the environment.
Cost and decision
The AI Cube Pro costs EUR 5,999 excl. VAT as a one-time purchase. It includes hardware, complete base setup and hardening, Open WebUI, model runtime, a local model agreed in advance, documentation and five hours of integration and configuration.
The comparison with cloud APIs depends on the real usage profile. For low or sporadic use, an API may be more economical. For sustained use, sensitive data or required technical control, owned hardware may be more suitable.
View the AI Cube and complete delivery scope or discuss your workload.
Related articles
Sources
Configure an AI Cube for your local AI workload
We translate models, data volume and usage patterns into a suitable configuration and deliver the AI Cube fully configured and hardened.
Frequently Asked Questions
Answers to important questions about this topic
Local AI inference runs an already trained model on owned hardware or controlled private infrastructure. Prompts and responses do not need to be sent to a public model API.
The scope includes the GB10 appliance with 128 GB unified memory, complete base setup and hardening, Open WebUI, Ollama and/or vLLM, a local model agreed in advance, documentation and five hours of integration and configuration.
This depends on model, context length, answer length and concurrency. A fixed user count based only on hardware specifications would not be reliable, so the workload is assessed before the model is fixed.
Yes. Open WebUI includes knowledge features for initial RAG use cases. Larger or permission-sensitive collections may require separate RAG pipelines, vector databases and interfaces.
No. Local processing improves control over data paths. Purpose, data, permissions, retention, logging and organisational measures remain decisive.

Written by
Timo Wevelsiep
Co-Founder & CEO
Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.
LinkedInLet's Talk About Your Idea
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.





