Which LLMs run on 128 GB of unified memory?
Timo Wevelsiep•Updated: 15.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Select the right model rather than merely installing one? For the AI Cube Pro, we agree the local model for language, task, and expected use in advance. Installation and functional testing are included. Explore the AI Cube Pro
With 128 GB of unified memory, GB10 systems can run model classes that do not fit into a typical workstation GPU. The number alone does not determine which LLM runs well. Weights, quantisation, runtime, KV cache, and memory reserved for the system must fit together.
What “fits” means in production
A model only fits when more than its weights can be loaded. The operating system, inference server, activations, and KV cache need memory too. KV-cache demand grows with context and active sequences. A model that starts in a single-user test may be unsuitable for ten concurrent chats.
As a rough calculation, weights use about two bytes per parameter at 16 bit, one byte at 8 bit, and half a byte at 4 bit. Real artefacts include metadata and quantisation structures, and the runtime needs headroom.
Specific model examples for GB10 and 128 GB
As of August 2026, NVIDIA's own DGX Spark playbooks list the following supported configurations among others. The list establishes technical support for those artefacts; it does not make each model suitable for every business workload.
| Model example | Officially documented configuration | Practical interpretation |
|---|---|---|
| Qwen3.6-35B-A3B | supported through LM Studio on DGX Spark | a comparatively compact MoE starting point with more room for context and concurrent requests |
| Llama 3.3 70B Instruct | supported in NVFP4 with TensorRT-LLM | a large dense model class; quality, language, and throughput still require workload testing |
| GPT-OSS-120B | supported in MXFP4 with TensorRT-LLM on one system | a large local model endpoint; context and concurrency need a dedicated load test |
| Nemotron 3 Super 120B-A12B | supported in NVFP4 with TensorRT-LLM | a large MoE model; assess licence, output quality, and runtime dependency before use |
Model names and support status change. A deployment therefore records the exact artefact, runtime, and release.
Useful model classes
Compact models
Models from one to the lower tens of billions leave substantial memory for context and concurrency. They often suit classification, extraction, short drafts, internal search, and narrowly defined assistants.
Mid-sized models
Models in the middle tens of billions often balance quality, multilingual capability, and latency. With suitable quantisation, 128 GB can retain room for several active requests.
Large and mixture-of-experts models
Quantised models in the hundreds-of-billions range can become technically possible. For mixture-of-experts models, total parameters and active parameters per token describe different things. Memory and compute do not follow the same number.
NVIDIA describes support for models up to 200 billion parameters on a DGX Spark and certain larger configurations across two systems. This is not a performance commitment for every model. Production requires testing the exact artefact.
Why a model family is not enough
“Qwen”, “Llama”, or “Mistral” names a family, not a reproducible configuration. Document the exact model ID and release, licence, quantisation and artefact author, runtime and version, typical and maximum context, quality tests, language, and throughput under realistic concurrency.
The broader guide Which LLM should you self-host? adds licensing, language, model-card, and evaluation criteria to this hardware perspective.
Selection for the AI Cube Pro
For the AI Cube Pro, we clarify the first use case before delivery. We then select, install, and test a local model. The goal is not the largest possible parameter count, but a platform that answers quickly enough and handles the intended content effectively.
Open WebUI can later present several local endpoints or deliberately approved cloud models. Model access can be constrained through users and groups. A model can change without rebuilding the user interface and knowledge spaces from scratch.
When two AI Cubes or Custom make sense
Two AI Cubes can increase aggregate throughput for independent requests or distribute a model. The second path needs appropriate software and ConnectX-7 configuration and is not identical to one computer with 256 GB of memory. For high concurrency, several large resident models, or committed availability, AI Cube Custom or a GPU server may be the better architecture.
Sources
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
It can with an appropriate quantisation, but parameter count is insufficient. Runtime, format, context, KV cache, and memory available to the system determine whether it runs reliably with useful concurrency.
NVIDIA states support for models up to 200 billion parameters for the GB10 platform. This is a technical upper range for suitable formats, not a promise of interactive multi-user performance for every 200B model.
No. For many business tasks, a smaller model tested for the purpose provides better latency and concurrency. Evaluate quality with your own tasks rather than inferring it from parameter count.
WZ-IT agrees a local model for language, task, and usage before delivery, installs it, and performs a functional test. The exact release is documented and can be changed later.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for confidentiality professions
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





