WZ-IT Logo
Qwen hosting

Qwen Hosting in Germany: Run Qwen Models on a Dedicated GPU

You want to use Qwen for chat, coding or documents without sending prompts to a third-party API. We run the right Qwen checkpoint on a dedicated GPU server in a German data centre.

Operated in a German data centreOperation under a DPAvLLM and OpenAI-compatible API

Companies worldwide trust WZ-IT

Reviews

Companies worldwide trust WZ-IT

Stadtwerke BrühlDGHO e.V.ABCO Water SystemsGolem.deEVADXBnextGYMAInergyml&sOdiseo SolutionsAnnotaARGESweetConnect GmbH
Starting point

Why run Qwen yourself

Qwen is Alibaba's open-weight model family. The current models cover general chat with image input (Qwen3.8-27B), large mixture-of-experts models (Qwen3.5-122B-A10B) and dedicated coding models (Qwen3-Coder-Next). The checkpoints covered here are licensed under Apache 2.0.

On your own server the weights run locally. Prompts, documents and source code are not sent to Alibaba or any other model API, and costs do not depend on the number of tokens.

  • Data stays in Germany

    The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.

  • Predictable costs

    A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.

  • Model version under control

    Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.

Model facts

Qwen checkpoints at a glance

Vendor specifications and weight file size per format, with source.

CheckpointReleasedParametersContext (vendor)Weights per format
Qwen/Qwen3.8-27B08/202627.8B (dense, with vision)262,144 tokens native
Qwen/Qwen3.5-122B-A10B02/2026122B total / 10B active (MoE)262,144 tokens native
Qwen/Qwen3-Coder-Next02/202680B total / 3B active (MoE)262,144 tokens native

GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.

Licence

Qwen: licence and commercial use

Licence
Apache License 2.0 (Licence text)
Commercial use
Yes, including products and services of companies of any size.
Obligations and thresholds
Include the licence text and copyright notices when redistributing and mark changes. No revenue or user threshold.
Checked
As of October 2026

Not offered

Qwen3.8-Flash-Next is licensed under the Qwen Community License 1.0. It requires companies running a model-as-a-service or AI work assistant business to obtain a separate licence from Qwen before any commercial use. We therefore do not offer this model as a hosted service. Licence text

Technical classification, not legal advice. The licence text of the deployed model version is authoritative.

Open LLM licences compared
Server tier

Minimum recommended server tier for Qwen

The recommendation follows from the weight size plus headroom for context and runtime.

ModelFormatWeightsMinimum recommendedNote
Qwen3.8-27BBF1655.6 GBManaged GPU Server 96
Qwen3.8-27BFP830.9 GBManaged GPU Server 96
Qwen3.8-27BNVFP421.9 GBManaged GPU Server 96The weights alone occupy almost all of the 24 GB, leaving no room for context and runtime.
Qwen3.5-122B-A10BFP8127.2 GBManaged GPU Server 192
Qwen3.5-122B-A10BBF16250.2 GBManaged GPU Server 384
Qwen3-Coder-NextFP880.4 GBManaged GPU Server 192On 96 GB, around 15 GB remain for context and runtime after the weights. That only supports short contexts and few parallel requests.
Qwen3-Coder-NextBF16159.4 GBManaged GPU Server 384

Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

The Managed GPU Server tiers

Managed GPU Server 24

1 × RTX PRO 4000 Blackwell, 24 GB GPU memory

€699 excl. VAT / month, cancellable monthly

€499 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 96

1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory

€1,799 excl. VAT / month, cancellable monthly

€999 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 192

2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory

Price and term on request

View configuration

Relevant for this family

Managed GPU Server 384

4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory

Price and term on request

View configuration

The Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.

vLLM

Running Qwen with vLLM

Architecture, quantisation and distribution across several GPUs.

Architectures

Qwen3.8-27B uses the Qwen3_5ForConditionalGeneration architecture, Qwen3.5-122B-A10B its MoE variant and Qwen3-Coder-Next Qwen3NextForCausalLM. All three are listed among the models supported by vLLM.

Quantisation

Qwen publishes official FP8 checkpoints. The NVFP4 version of Qwen3.8-27B comes from NVIDIA and uses mixed precision: MLP layers in NVFP4, attention in FP8. We compare quality and memory use per format on the target hardware.

Tensor parallelism

Qwen3.5-122B-A10B has 32 attention heads, Qwen3-Coder-Next has 16. Both can be split across two or four GPUs in vLLM. Three GPUs do not divide the heads evenly, which is why the three-GPU server is intended for several models side by side.

Context

The models support 262,144 tokens natively. We set the configured context length according to available GPU memory and parallelism. The full vendor context is not automatically part of every tier.

Operated by WZ-IT

What operation includes

The Managed GPU Server is an operated model environment, not an empty server.

Model and vLLM

An agreed model in the agreed quantisation, served through a managed vLLM inference layer.

OpenAI-compatible API

Applications and coding clients connect to the server via base URL and API key.

Open WebUI

A managed chat interface for teams using the model without their own application.

24/7 monitoring

Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.

Updates with CVE assessment

Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.

DPA and documentation

Data processing agreement, documented configuration and a personal point of contact.

Scope, prices and multi-GPU servers are on the product page.

View Managed GPU Server

Model assessment

Have your model and server tier assessed

Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.

Which server tier should be assessed?

The selection is non-binding. We confirm model fit, licence and term before a proposal.

How should we get back to you?

We usually respond within one business day. Please do not send access credentials yet.

Frequently asked questions about Qwen hosting

Licence, server tier, data path and operation

For general chat, documents and image input, Qwen3.8-27B is the natural starting point. The minimum recommended tier is the Managed GPU Server 96, in BF16, FP8 or NVFP4. For software development with agents, Qwen3-Coder-Next is the coding option, with the Managed GPU Server 192 as the minimum recommended tier.

Qwen3.8-27B, Qwen3.5-122B-A10B and Qwen3-Coder-Next are licensed under Apache 2.0. Commercial use is permitted without a revenue or user threshold. Qwen3.8-Flash-Next is the exception: the Qwen Community License requires a separate licence for certain business models, so we do not offer that model.

No. With hosting, the model weights run on your dedicated server in a German data centre. There is no connection to an Alibaba API. Requests, documents and answers stay on the server and in the systems you connect.

In mixture-of-experts models all experts must reside in GPU memory, even though only some of them compute per token. The FP8 weights occupy around 80 GB. On 96 GB little room remains for context, KV cache and runtime, so we recommend at least two GPUs.

The model supports this context, but every long context occupies GPU memory for the KV cache. We set the context length to suit the server tier and concurrent requests and state the value in the proposal.

Before switching we check new checkpoints for licence, vLLM support and memory requirements. The switch takes place in an agreed maintenance window, and the previous version stays available until acceptance.

Yes. vLLM provides an OpenAI-compatible API. Applications that already use an OpenAI-compatible interface connect to the server via base URL and API key. Open WebUI is also available as a chat interface.

The entry point for Qwen is the Managed GPU Server 96, the minimum recommended tier for Qwen3.8-27B: €1,799 excl. VAT / month plus €999 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. Larger variants need a multi-GPU server, priced on request. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.