WZ-IT Logo

GPT-OSS 120B on the AI Cube: Run OpenAI's Open-Weight Model Locally

Timo Wevelsiep
Timo Wevelsiep
Updated: 14.08.2026
#AI #OpenAI #GPTOSS #SelfHosting #LocalAI #AIServer #OnPremise #OpenWeights

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

GPT-OSS 120B on the AI Cube: Run OpenAI's Open-Weight Model Locally

Configure an AI Cube with GPT-OSS 120B - fully configured and hardened GB10 appliance, Open WebUI, an agreed local model, initial setup and support. Schedule an initial consultation

OpenAI released GPT-OSS 120B as an open-weight reasoning model on 5 August 2025. Its native MXFP4 quantisation allows OpenAI to position the model for a single GPU with 80 GB of memory. This also makes it relevant for compact systems with 128 GB unified memory.

The key point is that loading a model into memory is not the same as operating a suitable multi-user system. Context length, concurrent requests, runtime and required response time must fit together.

Table of contents

What defines GPT-OSS 120B

GPT-OSS 120B is a mixture-of-experts model. OpenAI states that 5.1 billion parameters are active per token, reducing compute compared with a dense model of the same total size.

Property GPT-OSS 120B
Model type Text reasoning model with mixture of experts
Active parameters per token 5.1 billion
Weight format MXFP4
Memory class stated by OpenAI One 80 GB GPU
Licence Apache 2.0
Interfaces Designed in part for Responses API-compatible workflows

Open weight means the weights can be run and adapted locally. It does not mean that the full training dataset and training process are public.

Does GPT-OSS 120B fit the AI Cube?

The WZ-IT AI Cube uses NVIDIA GB10 with 128 GB unified memory shared between CPU and GPU. The MXFP4 weights of GPT-OSS 120B therefore fit the available memory class in principle.

For a useful deployment, we also assess:

  • required context length,
  • expected concurrent requests,
  • typical answer length,
  • RAG context and tool schemas,
  • support in the selected runtime,
  • required response time.

Longer context and more concurrent requests consume additional memory beyond the weights. We therefore do not promise a fixed user count based only on the 128 GB specification.

What is configured before delivery

The local model is not selected on the fly at the customer site. Before ordering, we discuss the workload and agree whether GPT-OSS 120B or a smaller model is more appropriate.

The AI Cube Pro delivery includes:

  • complete initial setup of Linux, GPU drivers and container runtime,
  • hardening of the base installation and preparation for operation,
  • fully configured Open WebUI,
  • Ollama and/or vLLM selected for the agreed use case,
  • the agreed local model installed and technically verified,
  • documentation of the delivered configuration,
  • initial setup, onboarding and support.

The AI Cube therefore arrives prepared for plug-and-play use. Company-specific data sources, SSO and applications are then connected in a controlled way in the target environment.

Open WebUI, Ollama or vLLM

Open WebUI is the employee-facing workspace. It combines chat, model selection, files, knowledge collections and permissions. A runtime performs model inference:

  • Ollama is useful for straightforward model management and moderate usage.
  • vLLM provides an OpenAI-compatible API and targets efficient inference and concurrent requests.

The right runtime depends on more than the model name. We consider Arm and GB10 support, model format, throughput and integration requirements.

RAG and company knowledge

GPT-OSS 120B can act as the generation model in a RAG pipeline. A reliable knowledge system needs more than a large model:

  1. Documents must be extracted, structured and kept current.
  2. Permissions from source systems must be preserved.
  3. Embeddings and retrieval determine which content reaches the model.
  4. Answers should expose citations or source passages.
  5. Quality must be tested with real questions and expected answers.

Open WebUI includes knowledge features for initial use cases. For large or permission-sensitive collections, we can build custom RAG pipelines, middleware and interfaces.

Limits of the GB10 class

One AI Cube is compact and memory-rich, but it is not an HA GPU cluster. A larger architecture makes sense for:

  • many concurrent users with committed response times,
  • high availability and defined recovery objectives,
  • large models with long contexts,
  • high batch throughput,
  • training or substantial fine-tuning,
  • rack and data centre requirements.

Two AI Cubes can be linked through ConnectX-7. Whether a workload benefits depends on the runtime and parallelisation method.

Decision

GPT-OSS 120B is a plausible option for the AI Cube when a capable local reasoning model is required and the actual workload fits the GB10 class. For some tasks, a smaller model delivers better quality, latency and economics. This is why the model is agreed before delivery rather than fixed for every customer.

View the AI Cube and complete delivery scope or discuss the workload.

Sources

Enquiry

Run GPT-OSS 120B locally

We assess context, concurrency and quality requirements, agree the model before delivery and fully configure the AI Cube for the intended workload.

What are you planning with GPT-OSS 120B?

How should we get back to you?

Frequently Asked Questions

Answers to important questions about this topic

The AI Cube's GB10 hardware provides 128 GB unified memory and can hold an appropriate GPT-OSS 120B deployment. Whether latency and concurrent throughput fit the use case depends on runtime, quantisation, context and user count and is assessed before the model is fixed.

OpenAI describes the MXFP4 variant for a single 80 GB GPU. Context, runtime and concurrent requests require additional memory. The GB10 class has 128 GB unified memory, but workload testing is still necessary.

GPT-OSS is an open-weight model under the Apache 2.0 licence. Open model weights do not mean that the complete training dataset or training process is open.

If GPT-OSS fits after the initial assessment, we install and test it before delivery. The AI Cube arrives with a fully configured and hardened base, Open WebUI and the agreed model runtime.

Yes, it can be used as the generation model in a RAG architecture. Quality and permissions also depend on document preparation, embeddings, retrieval, source display and access design.

Timo Wevelsiep

Written by

Timo Wevelsiep

Co-Founder & CEO

Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.

LinkedIn

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • Maho Management
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.