25.06.2026
Local AI for Law Firms: §203, Case Data and RAG
How local RAG systems unlock case knowledge, cite sources and preserve matter-level permissions in every answer.
The AI Cube Standard is the plug-and-play entry point for AI chat, document search and custom assistants. Fully configured with an agreed local model, initial setup and support.

AI Cube Standard · Personal delivery and on-site setup available in Germany
The AI Cube provides local AI as a working tool for confidential documents, internal knowledge and recurring tasks across industries.
Search manuals, policies, files or project documentation and receive answers with source references.
Summarise, classify and compare documents and prepare them for downstream work.
Provide custom assistants with a defined task, selected model and approved knowledge base.
Prepare drafts, analyses and recurring text-based work with local models.
Process audio, dictation, images and PDFs locally depending on the model in use.
Provide AI to employees without requiring every user to have an account with an external model provider.
You receive a complete local AI platform that lets your business test its own use cases and run initial workloads in production. WZ-IT prepares the hardware, platform and model together.
AI Cube Standard
Fully configured local AI server for AI chat, document search, custom assistants and further applications inside the company network.
excl. VAT, one-time · optional managed operations from EUR 149.90 excl. VAT per month
Before delivery, we clarify the use case, model, storage and access paths. We then configure and test the system in Germany. The five included hours are available for commissioning, network and user configuration, onboarding and technical questions.
Compact appliance
Unified memory for local models and long contexts
The exact SSD and model configuration is defined before ordering based on data volume, model size and expected usage.
Confidential data and professional secrecy
With a fully local configuration, prompts, documents and model responses remain inside your defined infrastructure. This matters to businesses with confidential knowledge and to legal, healthcare, tax advisory and other confidentiality-sensitive environments.
For confidentiality-sensitive use, we align data paths, permissions, logging and administrative access. The legal and organisational assessment still depends on the specific use case.
Prompts, documents and model responses can remain within your own infrastructure.
Access to models and knowledge bases can be structured around people and teams.
Personal and shared knowledge bases receive appropriate read and write access.
Network paths, external providers, updates and administrative access are defined explicitly.
Chat in the browser, query documents and use custom assistants: Open WebUI is the shared workspace while models, knowledge and access remain under your control.
Preinstalled
Open WebUI combines chats, files, knowledge bases, models and tools. WZ-IT installs the platform on the Cube; authorised users can then create their own knowledge bases and share them with teams.
Chats, histories, files and model selection in an interface that teams can use without specialist tooling.
Authorised users can upload documents, create collections and use knowledge personally or together as a team.
Ollama or vLLM serves selected open-weight models directly on the AI Cube.
OIDC sign-in and appropriate access structures can be integrated with your identity environment.
Approved cloud or in-region models can be added selectively alongside local models.
In-house applications and reviewed tools can be connected through controlled interfaces.
WZ-IT connects users, networking, models and approved data sources. Your team works inside the company network or, where needed, through controlled remote access.
Target architecture
The Cube can operate in an isolated internal network or connect deliberately to existing platforms and approved model services. External access is not implemented by exposing an administration interface directly, but through defined network and identity paths.


OIDC, groups and roles for employees and administrators.
Local inference plus approved external endpoints where required.
Documents, RAG storage and existing data sources.
Updates, backup, monitoring and documented administration.
Local models and data sources without outbound model calls, suited to particularly sensitive use cases.
Employees and administrators connect through an encrypted identity-based VPN without exposing the Cube directly to the public internet.
For external use, we can design an upstream gateway with TLS, access controls and clear network separation.
Which data stays local and which external services are reachable is defined and documented before commissioning.
From secure integration to a custom application, you get one partner for the AI platform, infrastructure, networking and software development.
End-to-end implementation
You do not need to coordinate hardware, AI platform, networking and custom development across multiple providers. WZ-IT can implement and operate the complete technical path.
The five included hours cover commissioning, user and network configuration, onboarding and technical questions. Filling knowledge bases, data migrations and custom integrations are quoted separately.
Design and implement networking, users, firewalling, VPN access, storage and backup around the existing environment.
Connect Nextcloud, file systems, databases, wikis or suitable interfaces of line-of-business systems as a separate integration project.
Develop APIs, automation, middleware or MCP servers when standard features do not cover the required process.
Monitoring, patch management and technical support from EUR 149.90 excl. VAT per month, plus model changes, new integrations and further development under an agreed scope.
Once the configuration is approved, we prepare the hardware, platform and local model, deliver the fully configured AI Cube in under two weeks and commission it with you.
Align the use case, local model, storage, users and access paths together.
Fully configure, harden and technically verify the system, drivers, Open WebUI, inference stack and the local model agreed in advance.
Within two weeks of approval, we deliver the preconfigured AI Cube and commission it in your environment during a joint remote appointment.
We explain how to use the platform and support setup. Managed operations, monitoring and patches are then available from EUR 149.90 excl. VAT per month.
The included time is available for commissioning, network and user configuration, onboarding and technical questions. For data migrations, knowledge sources, and custom integrations, we can provide a clearly defined implementation scope.
After commissioning, your team can operate the AI Cube itself or hand ongoing technical responsibility to WZ-IT. The hardware and root access remain yours.
Managed AI Cube
excl. VAT per AI Cube per month
The base service covers the agreed technical operation of the existing installation. Scope, maintenance windows, backup targets and responsibilities are documented before service starts; extensions and larger integrations remain separately plannable.
Technical monitoring of the agreed system components with documented alert routes.
Scheduled security and system updates for the operating system, containers and the agreed AI stack.
Backup management, checks and recovery procedures according to the agreed scope and storage target.
Essential is the base level. Business, Production and Critical add defined response times; Production and Critical cover round-the-clock P1 response.
Tell us briefly how you want to use local AI and how many people will work with it. We will propose a suitable configuration and next step.
The AI Cube Standard covers the initial rollout and first production workloads. When models, concurrency, storage or availability requirements exceed that scope, we design the appropriate larger solution.
AI Cube Standard
For businesses introducing local AI, testing use cases with their own data and running first workloads in production.

AI Cube Custom
For larger models, more concurrent users, additional storage capacity, redundancy or specific network and operational requirements.

Local chat, document search and initial RAG applications for one team
AI Cube Standard
One site, 128 GB unified memory and predictable usage
AI Cube Standard
High concurrency, HA, rackmount or dedicated GPUs
AI Cube Custom
Existing GPU hardware should remain in use
Integration or Managed AI
No AI hardware wanted at your own site
LLM Hosting
Connect multiple systems
Two AI Cubes can be linked directly through a 200 GbE QSFP connection. Supported inference and fine-tuning workloads can then be distributed across nodes, including models that do not fit efficiently on a single system.
Connecting systems does not automatically provide high availability or improve every workload. Software, model, parallelisation method and network topology must fit the workload; any HA commitment is designed separately.
NVIDIA GB10 class with 128 GB unified memory and a choice of 1 TB, 2 TB or 4 TB NVMe SSD. Compact appliance form factor, ConnectX-7 for linking several Cubes, 230 V operation. Fully configured and hardened with Linux, GPU drivers, Open WebUI, model runtime and the local model agreed in advance. Root access is yours.
Topics
The AI Cube is a preconfigured local AI appliance for businesses. It combines Open WebUI, local models and internal knowledge sources for document search, internal assistants, transcription and technical AI workloads.
The WZ-IT AI Cube is our standard product based on validated ASUS/NVIDIA appliance hardware. For larger models, many concurrent users or special infrastructure requirements, we design AI Cube Custom builds with dedicated NVIDIA GPUs, multi-GPU, rackmount or custom networking.
Power consumption depends on the appliance or custom setup and the actual workload. The standard AI Cube is designed as a compact local AI box for office and enterprise environments; larger custom systems are assessed for power, cooling and site requirements in advance.
On the standard AI Cube the M.2 NVMe storage is replaceable. GPU and unified memory are part of the GB10 superchip and therefore fixed - more compute or more memory comes via an AI Cube Custom build, where we pick components freely. Either way the hardware is yours, so you can service it yourself or have it serviced.
It can, if only local models and local data sources are configured. Once external model providers, tools or integrations are enabled, data may leave the environment. We define and document these data paths during setup.
The AI Cube provides a technical foundation for local processing but does not replace a legal and organisational assessment. Model providers, data sources, permissions, logging, retention and administrative access must match the specific use case and be documented.
The AI Cube is technically suited to confidentiality-sensitive environments because models and data can be processed locally and access paths can be controlled. Whether a specific deployment meets legal and organisational requirements must be assessed for the individual use case.
Yes. Once the configuration is approved, we deliver the fully prepared AI Cube in under two weeks and commission it with you during an agreed remote appointment within that period. Hardware, Open WebUI, the model runtime and the agreed local model are configured, hardened and technically tested.
Yes. We handle the complete initial setup of the operating system, GPU drivers, container runtime, Open WebUI and model runtime, harden the base installation and verify the agreed local model before delivery.
The scope includes hardware, complete initial setup and hardening, the configured AI stack, a local model agreed and tested in advance, documentation and five hours for commissioning, network and user configuration, onboarding and technical questions. Populating knowledge bases, data migrations and custom integrations are quoted separately.
The AI Cube is prepared with Open WebUI as the interface and Ollama and/or vLLM for local inference. Open WebUI already includes knowledge collections and RAG. Local open-weight models, approved external models and additional vector databases are configured to match the use case.
Yes. Authorized users can upload documents in Open WebUI, create personal or shared knowledge bases and grant teams read or write access. Populating existing repositories, data migrations and automated integrations with DMS, Nextcloud or SharePoint are separate integration services.
Yes - depending on hardware configuration, multiple models can run in parallel. For intensive or parallel use, we recommend more powerful or customized hardware configurations.
Beyond chatbots and RAG systems: audio/video transcription, document indexing, data processing, code assistance, automation of internal processes - ideal for privacy-critical or compliance-relevant scenarios.
The standard AI Cube costs EUR 5,990.90 excl. VAT including hardware, complete base setup and hardening, an agreed local model and five hours for commissioning, onboarding and support. AI Cube Custom builds are quoted per project. The AI Cube becomes attractive especially for sensitive data, predictable workloads and long-term use: no external token dependency, full control over data and hardware.
When data privacy, control, consistent performance, and long-term planning are important - e.g., with sensitive data, compliance requirements, or frequent AI use.
Yes. We support migration: data and model transfer, re-setup on your on-prem system - without external dependency.
The AI Cube is purchased and owned by your company, while our AI servers are rented and run as a monthly managed service. The Cube is suitable for long-term planning, local control and fixed sites; rented servers are better for flexible projects or variable workloads.
The operating system, GPU drivers, containers, Open WebUI, inference stack and models require regular updates and an appropriate backup strategy. Your team can handle this or engage WZ-IT from EUR 149.90 excl. VAT per AI Cube per month for the agreed operations, monitoring and patch management.
Yes. We connect the Cube to existing networks, identity services and data sources. For controlled external access, we can use NetBird or an upstream reverse proxy gateway without exposing the administration interface directly to the public internet.
Managed operations start at EUR 149.90 excl. VAT per AI Cube per month. Depending on the agreed scope, this covers monitoring, patch management, backup coordination and technical support. Business, Production and Critical service levels are available for time-sensitive systems.
On request, we provide a backup concept: regular snapshots, redundant or external storage options, remote backup - keeping you protected even in case of hardware failure.
We provide remote commissioning and support in German and English. On-site appointments or partner delivery can be planned for suitable projects.
The standard AI Cube is based on validated ASUS/NVIDIA appliance hardware. Configuration, installation of the local AI stack, testing and commissioning are done by us in Germany. For AI Cube Custom builds we additionally select components to match the workload and integrate them at our site.
Yes - we offer a reseller program with attractive purchasing conditions, technical support, and optional white-label license. Ideal for system integrators and IT service providers.
25.06.2026
How local RAG systems unlock case knowledge, cite sources and preserve matter-level permissions in every answer.
10.05.2026
Three unauthenticated API calls. No login, no exploit framework, no privilege escalation. Three POST requests to a default port, and the machine's memory is on...
28.04.2026
DGX Spark and the WZ-IT AI Cube are not fundamentally different performance classes today. Both use NVIDIA's GB10 class with 128 GB unified memory. Comparing...
24.11.2025
OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a...
09.11.2025
Local AI inference means running a language or multimodal model on owned hardware rather than through a public API. Organisations gain control over data paths,...
08.11.2025
The use of Large Language Models (LLMs) such as GPT-4, Claude or Llama has evolved from experimental applications to mission-critical tools in recent years. However,...
No risk: worst case, you leave with a clearer understanding of your project than before.


“WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.”
From local AI integration to architecture, data sovereignty and ongoing operations.
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.
Timo Wevelsiep & Robin Zins
Managing Directors of WZ-IT

