WZ-IT Logo
gpt-oss hosting

gpt-oss Hosting in Germany: OpenAI's Open-Weight Models on a Dedicated GPU

You want to use a reasoning model from OpenAI without sending data to the OpenAI API. We run gpt-oss-120b or gpt-oss-20b on a dedicated GPU server in a German data centre.

Operated in a German data centreOperation under a DPAvLLM and OpenAI-compatible API

Companies worldwide trust WZ-IT

Reviews

Companies worldwide trust WZ-IT

Stadtwerke BrühlDGHO e.V.ABCO Water SystemsGolem.deEVADXBnextGYMAInergyml&sOdiseo SolutionsAnnotaARGESweetConnect GmbH
Starting point

Why run gpt-oss yourself

gpt-oss are OpenAI's open-weight models. Both variants are mixture-of-experts models with adjustable reasoning effort and tool calling. The expert weights were trained in MXFP4 format, so the checkpoints are small for their parameter count.

Running gpt-oss yourself means using the model without any connection to OpenAI. Costs depend on the server, not on consumed tokens, and data does not leave the German data centre through a model API.

  • Data stays in Germany

    The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.

  • Predictable costs

    A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.

  • Model version under control

    Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.

Model facts

gpt-oss checkpoints at a glance

Vendor specifications and weight file size per format, with source.

CheckpointReleasedParametersContext (vendor)Weights per format
openai/gpt-oss-120b08/2025117B total / 5.1B active (MoE)131,072 tokens
openai/gpt-oss-20b08/202521B total / 3.6B active (MoE)131,072 tokens

GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.

Licence

gpt-oss: licence and commercial use

Licence
Apache License 2.0 + gpt-oss Usage Policy (Licence text, Usage policy)
Commercial use
Yes, without a revenue or user threshold.
Obligations and thresholds
Include the licence text and notices when redistributing. The bundled usage policy requires compliance with applicable law.
Checked
As of October 2026

Technical classification, not legal advice. The licence text of the deployed model version is authoritative.

Open LLM licences compared
Server tier

Minimum recommended server tier for gpt-oss

The recommendation follows from the weight size plus headroom for context and runtime.

ModelFormatWeightsMinimum recommendedNote
gpt-oss-120bMXFP465.3 GBManaged GPU Server 96on premises: AI Cube (128 GB unified memory)For on-premises operation, the AI Cube with 128 GB unified memory is the alternative.
gpt-oss-20bMXFP413.8 GBManaged GPU Server 24

Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

The Managed GPU Server tiers

Relevant for this family

Managed GPU Server 24

1 × RTX PRO 4000 Blackwell, 24 GB GPU memory

€699 excl. VAT / month, cancellable monthly

€499 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 96

1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory

€1,799 excl. VAT / month, cancellable monthly

€999 excl. VAT one-time setup

View configuration

Managed GPU Server 192

2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory

Price and term on request

View configuration

Managed GPU Server 384

4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory

Price and term on request

View configuration

The Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.

vLLM

Running gpt-oss with vLLM

Architecture, quantisation and distribution across several GPUs.

Architecture

Both models use GptOssForCausalLM, an architecture supported by vLLM. gpt-oss-120b has 128 experts and gpt-oss-20b has 32, with 4 active per token in each.

MXFP4

The expert weights are stored in MXFP4, attention and embeddings in BF16. We determine which kernel runs on the RTX PRO 6000 Blackwell with the deployed vLLM version and verify it before the proposal.

Tensor parallelism

With 64 attention heads and 8 KV heads, gpt-oss can be split across two or four GPUs. A single 96 GB GPU is usually sufficient for gpt-oss-120b; more GPUs add room for longer contexts and more parallel requests.

Reasoning and format

gpt-oss uses the harmony response format with a separate reasoning channel. vLLM exposes it through the OpenAI-compatible API. You choose the reasoning effort (low, medium, high) per request.

Operated by WZ-IT

What operation includes

The Managed GPU Server is an operated model environment, not an empty server.

Model and vLLM

An agreed model in the agreed quantisation, served through a managed vLLM inference layer.

OpenAI-compatible API

Applications and coding clients connect to the server via base URL and API key.

Open WebUI

A managed chat interface for teams using the model without their own application.

24/7 monitoring

Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.

Updates with CVE assessment

Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.

DPA and documentation

Data processing agreement, documented configuration and a personal point of contact.

Scope, prices and multi-GPU servers are on the product page.

View Managed GPU Server

Model assessment

Have your model and server tier assessed

Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.

Which server tier should be assessed?

The selection is non-binding. We confirm model fit, licence and term before a proposal.

How should we get back to you?

We usually respond within one business day. Please do not send access credentials yet.

Frequently asked questions about gpt-oss hosting

Licence, server tier, data path and operation

gpt-oss-120b is the larger model for demanding reasoning and agent tasks; the minimum recommended tier is the Managed GPU Server 96. gpt-oss-20b is intended for simpler tasks, classification and fast answers; the minimum recommended tier is the Managed GPU Server 24.

No. gpt-oss are separate open-weight models that OpenAI published for self-hosting. With hosting there is no connection to OpenAI. Features and answer quality do not automatically match the models in ChatGPT.

Yes. The models are licensed under Apache 2.0, supplemented by a short usage policy that requires compliance with applicable law. There is no revenue or user threshold.

In many cases, yes. vLLM provides an OpenAI-compatible API; applications switch via base URL, API key and model name. Features only offered by the OpenAI platform, such as hosted tools, are checked case by case.

Yes, the AI Cube with 128 GB unified memory is an alternative to the hosted server. Whether it covers your use case depends on context length and concurrent requests. We check this before the proposal.

gpt-oss supports up to 131,072 tokens. The configured context length depends on the server tier and parallelism and is stated in the proposal.

The entry point for gpt-oss is the Managed GPU Server 24, the minimum recommended tier for gpt-oss-20b: €699 excl. VAT / month plus €499 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. For on-premises operation inside your own network, there is the AI Cube. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.