WZ-IT Logo
Mistral hosting

Mistral Hosting in Germany: Run Mistral Models on a Dedicated GPU

You want to run models from the French vendor Mistral AI yourself and classify the licence of each model correctly. We run the right checkpoint on a dedicated GPU server in a German data centre.

Operated in a German data centreOperation under a DPAvLLM and OpenAI-compatible API

Companies worldwide trust WZ-IT

Reviews

Companies worldwide trust WZ-IT

Stadtwerke BrühlDGHO e.V.ABCO Water SystemsGolem.deEVADXBnextGYMAInergyml&sOdiseo SolutionsAnnotaARGESweetConnect GmbH
Starting point

Why run Mistral yourself

Mistral AI publishes several model lines under different licences. Mistral Small 4 combines chat, reasoning and coding in one mixture-of-experts model under Apache 2.0. Devstral Small 2 targets software development. Mistral Medium 3.5 and Devstral 2 are larger dense models under a Modified MIT License with a revenue threshold.

Because licences differ within the family, we check before the proposal which checkpoint is permitted for your company. The model runs on your dedicated server, without a connection to the Mistral API.

  • Data stays in Germany

    The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.

  • Predictable costs

    A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.

  • Model version under control

    Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.

Model facts

Mistral checkpoints at a glance

Vendor specifications and weight file size per format, with source.

CheckpointReleasedParametersContext (vendor)Weights per format
mistralai/Mistral-Small-4-119B-260303/2026119B total / 6.5B active (MoE)256K tokens
mistralai/Devstral-Small-2-24B-Instruct-251212/202524B (dense)256K tokens
mistralai/Mistral-Medium-3.5-128B05/2026128B (dense)256K tokens
mistralai/Devstral-2-123B-Instruct-251212/2025123B (dense)256K tokens

GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.

Licence

Mistral: licence and commercial use

Licence
Apache 2.0 (Small 4, Devstral Small 2) / Modified MIT (Medium 3.5, Devstral 2) (Licence text)
Commercial use
Mistral Small 4 and Devstral Small 2: yes, without a threshold. Mistral Medium 3.5 and Devstral 2: only for companies whose consolidated global monthly revenue did not exceed USD 20 million in the preceding month.
Obligations and thresholds
Modified MIT: include attribution and the licence notice. Above the revenue threshold the licence grants no rights, including for derived models; Mistral AI grants commercial licences on request.
Checked
As of October 2026

Technical classification, not legal advice. The licence text of the deployed model version is authoritative.

Open LLM licences compared
Server tier

Minimum recommended server tier for Mistral

The recommendation follows from the weight size plus headroom for context and runtime.

ModelFormatWeightsMinimum recommendedNote
Mistral Small 4NVFP470.8 GBManaged GPU Server 96
Mistral Small 4FP8120.9 GBManaged GPU Server 192
Devstral Small 2FP825.8 GBManaged GPU Server 96
Mistral Medium 3.5FP8133.6 GBManaged GPU Server 192Licence note: no rights above USD 20 million consolidated monthly revenue.
Devstral 2FP8128.2 GBManaged GPU Server 192Licence note: no rights above USD 20 million consolidated monthly revenue.

Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

The FP8 sizes are the sum of the weight files of one format. The Mistral repositories also contain the weights a second time in consolidated format, which is not loaded twice in operation.

The Managed GPU Server tiers

Managed GPU Server 24

1 × RTX PRO 4000 Blackwell, 24 GB GPU memory

€699 excl. VAT / month, cancellable monthly

€499 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 96

1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory

€1,799 excl. VAT / month, cancellable monthly

€999 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 192

2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory

Price and term on request

View configuration

Managed GPU Server 384

4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory

Price and term on request

View configuration

The Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.

vLLM

Running Mistral with vLLM

Architecture, quantisation and distribution across several GPUs.

Architecture

Mistral Small 4, Mistral Medium 3.5 and Devstral Small 2 use Mistral3ForConditionalGeneration, an architecture supported by vLLM. Mistral recommends a current version of the mistral_common package for serving.

NVFP4 and FP8

Mistral officially publishes Mistral Small 4 in FP8 and NVFP4. The NVFP4 version occupies around 71 GB and is therefore the variant for the Managed GPU Server 96. In the model card, Mistral uses two GPUs for both formats.

Attention

Mistral Small 4 uses multi-head latent attention. vLLM needs a suitable attention backend for it, which we verify with the deployed version on the RTX PRO 6000 Blackwell.

Tensor parallelism

Mistral Medium 3.5 and Devstral 2 have 96 attention heads and 8 KV heads and can be split across two or four GPUs. The vendor examples use eight GPUs; for the FP8 weights of around 130 GB we recommend at least two RTX PRO 6000.

Operated by WZ-IT

What operation includes

The Managed GPU Server is an operated model environment, not an empty server.

Model and vLLM

An agreed model in the agreed quantisation, served through a managed vLLM inference layer.

OpenAI-compatible API

Applications and coding clients connect to the server via base URL and API key.

Open WebUI

A managed chat interface for teams using the model without their own application.

24/7 monitoring

Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.

Updates with CVE assessment

Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.

DPA and documentation

Data processing agreement, documented configuration and a personal point of contact.

Scope, prices and multi-GPU servers are on the product page.

View Managed GPU Server

Model assessment

Have your model and server tier assessed

Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.

Which server tier should be assessed?

The selection is non-binding. We confirm model fit, licence and term before a proposal.

How should we get back to you?

We usually respond within one business day. Please do not send access credentials yet.

Frequently asked questions about Mistral hosting

Licence, server tier, data path and operation

Mistral Small 4 in NVFP4 and Devstral Small 2 in FP8. Both are licensed under Apache 2.0. Mistral Small 4 in FP8, Mistral Medium 3.5 and Devstral 2 need at least the Managed GPU Server 192.

Only if your company's consolidated global monthly revenue did not exceed USD 20 million in the preceding month. Above that the Modified MIT License grants no rights; Mistral AI grants commercial licences on request. Mistral Small 4 and Devstral Small 2 are under Apache 2.0 without a threshold.

The origin of the model does not determine data protection. What matters is where and how it runs. With us, the model runs on a dedicated server in a German data centre, under a data processing agreement and without a connection to the Mistral API.

Devstral Small 2 is a dense model with 24B parameters and runs on the Managed GPU Server 96. Qwen3-Coder-Next is larger and needs at least the Managed GPU Server 192. We compare which model gives better results on your repositories using your tasks.

NVFP4 stores the weights mostly in 4 bits and occupies around 71 GB; FP8 occupies around 121 GB. For NVFP4 the minimum recommended tier is the Managed GPU Server 96, for FP8 the Managed GPU Server 192. We check quality differences using examples from your use case.

Yes. vLLM provides an OpenAI-compatible API, and Open WebUI is available as a chat interface. Coding clients with an OpenAI-compatible interface connect directly.

The entry point for Mistral is the Managed GPU Server 96, the minimum recommended tier for Mistral Small 4 and Devstral Small 2: €1,799 excl. VAT / month plus €999 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. Larger variants need a multi-GPU server, priced on request. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.