GPT-OSS 120B on the AI Cube: Run OpenAI's Open-Weight Model Locally

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

Configure an AI Cube with GPT-OSS 120B - fully configured and hardened GB10 appliance, Open WebUI, an agreed local model, initial setup and support. Schedule an initial consultation
OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a single GPU with 80 GB of memory. This also makes it relevant for compact systems with 128 GB unified memory.
The key point is that loading a model into memory is not the same as operating a suitable multi-user system. Context length, concurrent requests, runtime and required response time must fit together.
Table of contents
- What defines GPT-OSS 120B
- Does GPT-OSS 120B fit the AI Cube?
- What is configured before delivery
- Open WebUI, Ollama or vLLM
- RAG and company knowledge
- Limits of the GB10 class
- Decision
What defines GPT-OSS 120B
GPT-OSS 120B is a mixture-of-experts model. OpenAI states that 5.1 billion parameters are active per token, reducing compute compared with a dense model of the same total size.
| Property | GPT-OSS 120B |
|---|---|
| Model type | Text reasoning model with mixture of experts |
| Active parameters per token | 5.1 billion |
| Weight format | MXFP4 |
| Memory class stated by OpenAI | One 80 GB GPU |
| Licence | Apache 2.0 |
| Interfaces | Designed in part for Responses API-compatible workflows |
Open weight means the weights can be run and adapted locally. It does not mean that the full training dataset and training process are public.
Does GPT-OSS 120B fit the AI Cube?
The WZ-IT AI Cube uses NVIDIA GB10 with 128 GB unified memory shared between CPU and GPU. The MXFP4 weights of GPT-OSS 120B therefore fit the available memory class in principle.
For a useful deployment, we also assess:
- required context length,
- expected concurrent requests,
- typical answer length,
- RAG context and tool schemas,
- support in the selected runtime,
- required response time.
Longer context and more concurrent requests consume additional memory beyond the weights. We therefore do not promise a fixed user count based only on the 128 GB specification.
What is configured before delivery
The local model is not selected on the fly at the customer site. Before ordering, we discuss the workload and agree whether GPT-OSS 120B or a smaller model is more appropriate.
The AI Cube Pro delivery includes:
- complete initial setup of Linux, GPU drivers and container runtime,
- hardening of the base installation and preparation for operation,
- fully configured Open WebUI,
- Ollama and/or vLLM selected for the agreed use case,
- the agreed local model installed and technically verified,
- documentation of the delivered configuration,
- initial setup, onboarding and support.
The AI Cube therefore arrives prepared for plug-and-play use. Company-specific data sources, SSO and applications are then connected in a controlled way in the target environment.
Open WebUI, Ollama or vLLM
Open WebUI is the employee-facing workspace. It combines chat, model selection, files, knowledge collections and permissions. A runtime performs model inference:
- Ollama is useful for straightforward model management and moderate usage.
- vLLM provides an OpenAI-compatible API and targets efficient inference and concurrent requests.
The right runtime depends on more than the model name. We consider Arm and GB10 support, model format, throughput and integration requirements.
RAG and company knowledge
GPT-OSS 120B can act as the generation model in a RAG pipeline. A reliable knowledge system needs more than a large model:
- Documents must be extracted, structured and kept current.
- Permissions from source systems must be preserved.
- Embeddings and retrieval determine which content reaches the model.
- Answers should expose citations or source passages.
- Quality must be tested with real questions and expected answers.
Open WebUI includes knowledge features for initial use cases. For large or permission-sensitive collections, we can build custom RAG pipelines, middleware and interfaces.
Limits of the GB10 class
One AI Cube is compact and memory-rich, but it is not an HA GPU cluster. A larger architecture makes sense for:
- many concurrent users with committed response times,
- high availability and defined recovery objectives,
- large models with long contexts,
- high batch throughput,
- training or substantial fine-tuning,
- rack and data centre requirements.
Two AI Cubes can be linked through ConnectX-7. Whether a workload benefits depends on the runtime and parallelisation method.
Decision
GPT-OSS 120B is a plausible option for the AI Cube when a capable local reasoning model is required and the actual workload fits the GB10 class. For some tasks, a smaller model delivers better quality, latency and economics. This is why the model is agreed before delivery rather than fixed for every customer.
View the AI Cube and complete delivery scope or discuss the workload.
Related articles
Sources
Run GPT-OSS 120B locally
We assess context, concurrency and quality requirements, agree the model before delivery and fully configure the AI Cube for the intended workload.
Frequently Asked Questions
Answers to important questions about this topic
The AI Cube's GB10 hardware provides 128 GB unified memory and can hold an appropriate GPT-OSS 120B deployment. Whether latency and concurrent throughput fit the use case depends on runtime, quantisation, context and user count and is assessed before the model is fixed.
OpenAI describes the MXFP4 variant for a single 80 GB GPU. Context, runtime and concurrent requests require additional memory. The GB10 class has 128 GB unified memory, but workload testing is still necessary.
GPT-OSS is an open-weight model under the Apache 2.0 licence. Open model weights do not mean that the complete training dataset or training process is open.
If GPT-OSS fits after the initial assessment, we install and test it before delivery. The AI Cube arrives with a fully configured and hardened base, Open WebUI and the agreed model runtime.
Yes, it can be used as the generation model in a RAG architecture. Quality and permissions also depend on document preparation, embeddings, retrieval, source display and access design.

Written by
Timo Wevelsiep
Co-Founder & CEO
Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.
LinkedInLet's Talk About Your Idea
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.





