Configure NVIDIA DGX Spark for business
Timo Wevelsiep•Updated: 15.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Receive DGX Spark as an operational AI platform? WZ-IT configures the hardware, Open WebUI, local model runtime, and an agreed model, and can take responsibility for ongoing operations. Explore the AI Cube Pro
NVIDIA DGX Spark is the reference system for the GB10 platform. It combines a 20-core Arm processor, Blackwell GPU, and 128 GB of coherent unified memory in a compact computer. This accommodates local model classes that are impractical on ordinary workstation GPUs. A production AI service still requires a suitable software, security, and operating architecture.
Hardware and intended use
NVIDIA specifies up to 1 PFLOP FP4, a 4 TB self-encrypting NVMe SSD, 10 GbE, and ConnectX-7. DGX OS and NVIDIA's AI stack are preinstalled. Its ARM64 architecture must be considered for containers, Python dependencies, and custom extensions.
The phrase “up to 200 billion parameters” is not a performance guarantee. A model may fit in memory yet remain too slow with long context or several concurrent users. Selection therefore needs measurements with representative input.
From developer system to business platform
A production setup includes:
- documenting and updating firmware, DGX OS, drivers, and container runtime;
- named administrators, SSH keys, and least privilege;
- defined network, DNS, TLS, firewall, user, and administration paths;
- a model runtime and model selected for language, quality, and throughput;
- Open WebUI with users, groups, model permissions, and persistent storage;
- backup coverage for models, chat data, knowledge spaces, and configuration;
- monitoring of the system, services, storage, and model endpoint;
- documented updates, restore procedures, and support access.
What staff use afterwards
Users work through a familiar browser interface. Depending on permissions, they can choose approved local models, work with files, and maintain personal or shared knowledge spaces. Approved cloud models can be added without replacing the local platform.
Automated connections to Nextcloud, SharePoint, a DMS, or business applications are separate integration projects, as are custom RAG pipelines, APIs, and MCP servers.
When DGX Spark is the right choice
The reference system is a good fit when an organisation explicitly wants NVIDIA's DGX platform and a clearly documented vendor path. It may be excessive for one workstation without central users or operating requirements. One device may be insufficient for high concurrency, redundant service, or very large models.
Application, data volume, expected concurrency, and support model should therefore be defined before procurement. This determines whether one DGX Spark as an AI Cube is sufficient, a second node is useful, or a larger GPU platform is the better starting point.
AI Cube Pro in the DGX Spark class
The AI Cube Pro is WZ-IT's complete starter package in this hardware class. It includes hardware, a preconfigured AI platform, agreed local model, hardening, functional testing, initial setup, and five hours of support. It is generally ready within two weeks after configuration approval.
Additional AI Cubes can serve more users or larger models. Connecting AI Cubes with ConnectX-7 explains request balancing and distributed inference.
Sources
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
GB10 and 128 GB of unified memory provide a compact foundation for local inference, prototyping, and initial production AI applications. Suitability still depends on the model, usage pattern, and availability requirements.
DGX OS and NVIDIA's stack are preinstalled. Business use still requires a model runtime, user interface, access controls, hardening, backup, monitoring, and operating documentation for the intended application.
NVIDIA states a technical limit of models up to 200 billion parameters. In practice, quantisation, context, KV cache, runtime, and concurrent users determine a useful model choice.
Yes. Two systems can be connected directly through ConnectX-7. Depending on the objective, requests are distributed or a supported model is executed across both nodes.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for confidentiality professions
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





