Connect AI Cubes with ConnectX-7: plan a cluster
Timo Wevelsiep•Updated: 15.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Start with one AI Cube Pro and expand when required? The AI Cube Pro is designed as a productive single-system starting point. For more users or larger models, we connect and configure additional AI Cubes for the measured workload. Explore the AI Cube Pro and expansion paths
GB10 systems can connect directly through ConnectX-7. This enables two distinct architectures: multiple model endpoints for more concurrent requests, or one model distributed across nodes. Calling both a “cluster” hides different goals and bottlenecks.
Hardware foundation
Each GB10-based AI Cube has 128 GB of unified memory, 10 GbE, and two QSFP connectors attached to the ConnectX-7 NIC. NVIDIA documents up to 200 Gbit/s per QSFP port and recommends compatible cables. These interfaces operate as Ethernet; they do not create shared memory between systems.
Two devices can use a direct cable. The current NVIDIA Sync Cluster Assistant also supports three systems connected directly and up to four through a switch. Larger or different topologies need a custom network design. Cabling, firmware, MTU, addressing, and topology must be documented.
Option A: more independent requests
Each AI Cube loads a model and serves its own requests. A load balancer sends new sessions or API calls to available endpoints. This increases aggregate throughput, allows different models per node, and can isolate failures.
Open WebUI, database, knowledge store, and session state still need a defined location. Duplicated inference alone is not a highly available platform.
Option B: one model across systems
If a model and KV cache do not fit usefully on one node, a supported runtime can distribute weights and computation. Nodes exchange intermediate results over ConnectX-7. Latency and throughput then depend on model architecture, parallelisation, runtime, and network.
Two devices do not become one computer with 256 GB of freely addressable memory. Each provides 128 GB coordinated by distributed software, so the exact model artefact must be tested.
Direct two-node connection
A production setup includes compatible operating-system, driver, and firmware levels; a validated QSFP cable; explicit IP addresses; connectivity and throughput tests; runtime configuration; workload measurement; and monitoring for nodes, network, model, and queues.
NVIDIA notes that one physical port can appear as several Linux interfaces because of the internal topology. Use the documented port mapping rather than guessing from interface names.
What NVIDIA Sync handles - and what it does not
The NVIDIA Sync Cluster Assistant discovers devices, creates a supported ConnectX-7 topology, applies network settings, validates links, and configures SSH between nodes. This establishes the communication layer.
It does not configure model distribution, NCCL or vLLM, container orchestration, load balancing, or high availability. Those layers still need to be designed and tested for the model and operating objective. A successful network test is not yet a working inference cluster.
Decision matrix
| Goal | First approach |
|---|---|
| more concurrent users | independent model endpoints and load balancing |
| larger single model | distributed inference with exact model testing |
| several specialist models | separate models by node and route deliberately |
| committed availability | make the complete platform redundant |
| three nodes | directly cabled with NVIDIA Sync support; still validate topology and workload |
| four nodes | plan a supported switch topology and operating model |
| more than four nodes | custom network and orchestration design required |
WZ-IT expansion path
With the AI Cube Pro, one local workload is first configured and measured. As use grows, we add AI Cubes for throughput or distributed models. When rack mounting, redundant power, several discrete GPUs, or stricter availability are needed, AI Cube Custom or a GPU server is the better path.
Before scaling out, define both the usage and load profile and the specific model artefact for 128 GB. Otherwise, hardware is connected before the actual capacity problem has been identified.
Sources
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
Yes. GB10 systems provide ConnectX-7 interfaces that can link two devices with an appropriate QSFP cable. Production operation also needs correct IP, runtime, and workload configuration.
The current NVIDIA Sync Cluster Assistant supports up to three DGX Spark or GB10 systems connected directly and up to four systems through a switch. It configures networking and SSH, not distributed inference or workload orchestration.
No. They remain two systems with 128 GB of unified memory each. Distributed software divides a model or requests and transfers data over the network, which differs from shared addressable memory.
When independent requests are distributed across multiple local model endpoints through a load balancer. This increases aggregate throughput without distributing one model.
No. The interface, database, knowledge, identity, load balancer, and storage must also be redundant or recoverable, not only the model API.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for confidentiality professions
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





