Deployed worldwide
WZ-IT Logo

Which LLMs run on 128 GB of unified memory?

Timo WevelsiepTimo WevelsiepUpdated: 15.08.2026

Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.

Select the right model rather than merely installing one? For the AI Cube Pro, we agree the local model for language, task, and expected use in advance. Installation and functional testing are included. Explore the AI Cube Pro

With 128 GB of unified memory, GB10 systems can run model classes that do not fit into a typical workstation GPU. The number alone does not determine which LLM runs well. Weights, quantisation, runtime, KV cache, and memory reserved for the system must fit together.

What “fits” means in production

A model only fits when more than its weights can be loaded. The operating system, inference server, activations, and KV cache need memory too. KV-cache demand grows with context and active sequences. A model that starts in a single-user test may be unsuitable for ten concurrent chats.

As a rough calculation, weights use about two bytes per parameter at 16 bit, one byte at 8 bit, and half a byte at 4 bit. Real artefacts include metadata and quantisation structures, and the runtime needs headroom.

Specific model examples for GB10 and 128 GB

As of August 2026, NVIDIA's own DGX Spark playbooks list the following supported configurations among others. The list establishes technical support for those artefacts; it does not make each model suitable for every business workload.

Model example Officially documented configuration Practical interpretation
Qwen3.6-35B-A3B supported through LM Studio on DGX Spark a comparatively compact MoE starting point with more room for context and concurrent requests
Llama 3.3 70B Instruct supported in NVFP4 with TensorRT-LLM a large dense model class; quality, language, and throughput still require workload testing
GPT-OSS-120B supported in MXFP4 with TensorRT-LLM on one system a large local model endpoint; context and concurrency need a dedicated load test
Nemotron 3 Super 120B-A12B supported in NVFP4 with TensorRT-LLM a large MoE model; assess licence, output quality, and runtime dependency before use

Model names and support status change. A deployment therefore records the exact artefact, runtime, and release.

Useful model classes

Compact models

Models from one to the lower tens of billions leave substantial memory for context and concurrency. They often suit classification, extraction, short drafts, internal search, and narrowly defined assistants.

Mid-sized models

Models in the middle tens of billions often balance quality, multilingual capability, and latency. With suitable quantisation, 128 GB can retain room for several active requests.

Large and mixture-of-experts models

Quantised models in the hundreds-of-billions range can become technically possible. For mixture-of-experts models, total parameters and active parameters per token describe different things. Memory and compute do not follow the same number.

NVIDIA describes support for models up to 200 billion parameters on a DGX Spark and certain larger configurations across two systems. This is not a performance commitment for every model. Production requires testing the exact artefact.

Why a model family is not enough

“Qwen”, “Llama”, or “Mistral” names a family, not a reproducible configuration. Document the exact model ID and release, licence, quantisation and artefact author, runtime and version, typical and maximum context, quality tests, language, and throughput under realistic concurrency.

The broader guide Which LLM should you self-host? adds licensing, language, model-card, and evaluation criteria to this hardware perspective.

Selection for the AI Cube Pro

For the AI Cube Pro, we clarify the first use case before delivery. We then select, install, and test a local model. The goal is not the largest possible parameter count, but a platform that answers quickly enough and handles the intended content effectively.

Open WebUI can later present several local endpoints or deliberately approved cloud models. Model access can be constrained through users and groups. A model can change without rebuilding the user interface and knowledge spaces from scratch.

When two AI Cubes or Custom make sense

Two AI Cubes can increase aggregate throughput for independent requests or distribute a model. The second path needs appropriate software and ConnectX-7 configuration and is not identical to one computer with 256 GB of memory. For high concurrency, several large resident models, or committed availability, AI Cube Custom or a GPU server may be the better architecture.

Sources

Rather have it operated?

You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.

Enquiry

Assess local AI for your use case

Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.

How should we get back to you?

Frequently Asked Questions

Answers to the most important questions

It can with an appropriate quantisation, but parameter count is insufficient. Runtime, format, context, KV cache, and memory available to the system determine whether it runs reliably with useful concurrency.

NVIDIA states support for models up to 200 billion parameters for the GB10 platform. This is a technical upper range for suitable formats, not a promise of interactive multi-user performance for every 200B model.

No. For many business tasks, a smaller model tested for the purpose provides better latency and concurrency. Evaluate quality with your own tasks rather than inferring it from parameter count.

WZ-IT agrees a local model for language, task, and usage before delivery, installs it, and performs a functional test. The exact release is documented and can be changed later.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • Maho Management
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/2 - Topic Selection50%

What is your inquiry about?

Select one or more areas where we can support you.