Deployed worldwide
WZ-IT Logo

Local AI Inference with the AI Cube: Run AI Inside Your Own Network

Timo Wevelsiep
Timo Wevelsiep
Updated: 14.08.2026
#AI #SelfHosting #AIInference #DataProtection #AIServer #OnPremise #OpenWebUI

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

Local AI Inference with the AI Cube: Run AI Inside Your Own Network

View the WZ-IT AI Cube - local AI server with 128 GB unified memory, complete initial setup, hardening, Open WebUI and an agreed model. Schedule an initial consultation

Local AI inference means running a language or multimodal model on owned hardware rather than through a public API. Organisations gain control over data paths, model selection and cost structure. They also take responsibility for hardware, software, permissions, updates and operations.

The WZ-IT AI Cube combines these components in a compact local AI server. It is not delivered as an empty developer machine, but fully configured, hardened and prepared with a local model agreed in advance.

Table of contents

What is processed locally

In a fully local configuration, model, chat workspace, knowledge index and database run inside the organisation's own network. Prompts, uploaded files and model responses do not need to be transferred to an external model provider.

Whether there are truly no external connections depends on configuration. Cloud models, web search, external tools or update services create additional data paths. We define and document these connections during setup.

Local processing is particularly relevant when:

  • confidential documents are processed,
  • professional confidentiality or customer requirements apply,
  • public model APIs are not approved,
  • fixed local capacity is preferred over variable API usage,
  • custom integrations should avoid provider lock-in.

AI Cube hardware

The standard configuration uses a compact ASUS/NVIDIA appliance in the GB10 class:

Feature Standard configuration
Compute platform NVIDIA GB10 Grace Blackwell
Processor 20-core Arm
Memory 128 GB LPDDR5x unified memory
Storage 1, 2 or 4 TB NVMe depending on configuration
Networking 10 GbE and ConnectX-7
Form factor Compact desktop appliance

Unified memory is shared between CPU and GPU. It allows quantised models to fit into a memory class unavailable on many individual graphics cards. Actual speed still depends on model architecture, quantisation, context and concurrency.

What is configured before delivery

Before ordering, we discuss the workload and agree the local model. A larger model is not automatically better: for some tasks, a smaller model provides lower latency and more concurrency.

The AI Cube is then prepared for plug-and-play use:

  • Linux, GPU drivers and container runtime fully configured,
  • base installation hardened and prepared for operation,
  • Open WebUI installed and configured,
  • Ollama and/or vLLM set up for the workload,
  • agreed local model installed and verified,
  • delivered configuration documented.

Five hours for integration into your environment are also included. This time can cover network, identity, an initial knowledge source or model parameters. Larger integrations are planned separately.

Open WebUI as the workspace

Open WebUI provides a browser-based workspace. Employees receive a familiar chat interface while administrators control models, users, knowledge collections and connections.

Depending on configuration, it can provide:

  • chat and history with local models,
  • multiple approved models in one interface,
  • file uploads and knowledge collections,
  • roles, groups and OIDC integration,
  • custom assistants with system prompts and knowledge,
  • OpenAI-compatible interfaces for internal applications.

External cloud models can optionally be connected alongside local models. Sensitive workloads can also run with local models only.

RAG and internal documents

Retrieval-augmented generation adds relevant content from internal documents to a model. The process involves several stages:

  1. Documents are extracted and split into useful sections.
  2. Embeddings represent these sections for search.
  3. Relevant passages are retrieved for a question.
  4. The model generates an answer from that context.
  5. Citations make the answer traceable.

Open WebUI includes knowledge features for initial use cases. For large, frequently changing or permission-sensitive collections, we build separate pipelines, vector databases, middleware and interfaces to Nextcloud, BookStack, document management or line-of-business applications.

Sizing and limits

Reliable sizing considers at least:

  • model and quantisation,
  • prompt, context and answer length,
  • concurrent active users,
  • interactive or batch use,
  • RAG, image processing or tool calls,
  • availability and recovery requirements.

One AI Cube often fits local assistants, document search and inference for a limited team. For high concurrency, HA, rackmount or dedicated GPUs, we design custom GPU servers.

Two AI Cubes can be linked through ConnectX-7 and supported workloads can be distributed. Linking systems does not automatically create high availability or accelerate every model.

Self-operation or managed service

You receive root access and can operate the AI Cube yourself. Operations include:

  • security and operating system updates,
  • updates for Open WebUI and the model runtime,
  • model changes and compatibility checks,
  • monitoring of hardware, services and storage,
  • backups of configuration, database and knowledge collections,
  • controlled recovery.

WZ-IT can optionally take over these tasks. Scope and response times are agreed for the environment.

Cost and decision

The AI Cube Pro costs EUR 5,999 excl. VAT as a one-time purchase. It includes hardware, complete base setup and hardening, Open WebUI, model runtime, a local model agreed in advance, documentation and five hours of integration and configuration.

The comparison with cloud APIs depends on the real usage profile. For low or sporadic use, an API may be more economical. For sustained use, sensitive data or required technical control, owned hardware may be more suitable.

View the AI Cube and complete delivery scope or discuss your workload.

Sources

Enquiry

Configure an AI Cube for your local AI workload

We translate models, data volume and usage patterns into a suitable configuration and deliver the AI Cube fully configured and hardened.

Which next step are you planning?

How should we get back to you?

Frequently Asked Questions

Answers to important questions about this topic

Local AI inference runs an already trained model on owned hardware or controlled private infrastructure. Prompts and responses do not need to be sent to a public model API.

The scope includes the GB10 appliance with 128 GB unified memory, complete base setup and hardening, Open WebUI, Ollama and/or vLLM, a local model agreed in advance, documentation and five hours of integration and configuration.

This depends on model, context length, answer length and concurrency. A fixed user count based only on hardware specifications would not be reliable, so the workload is assessed before the model is fixed.

Yes. Open WebUI includes knowledge features for initial RAG use cases. Larger or permission-sensitive collections may require separate RAG pipelines, vector databases and interfaces.

No. Local processing improves control over data paths. Purpose, data, permissions, retention, logging and organisational measures remain decisive.

Timo Wevelsiep

Written by

Timo Wevelsiep

Co-Founder & CEO

Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.

LinkedIn

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • Maho Management
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/2 - Topic Selection50%

What is your inquiry about?

Select one or more areas where we can support you.