You want to run models from the French vendor Mistral AI yourself and classify the licence of each model correctly. We run the right checkpoint on a dedicated GPU server in a German data centre.
Companies worldwide trust WZ-IT
Mistral AI publishes several model lines under different licences. Mistral Small 4 combines chat, reasoning and coding in one mixture-of-experts model under Apache 2.0. Devstral Small 2 targets software development. Mistral Medium 3.5 and Devstral 2 are larger dense models under a Modified MIT License with a revenue threshold.
Because licences differ within the family, we check before the proposal which checkpoint is permitted for your company. The model runs on your dedicated server, without a connection to the Mistral API.
The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.
A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.
Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.
Vendor specifications and weight file size per format, with source.
| Checkpoint | Released | Parameters | Context (vendor) | Weights per format |
|---|---|---|---|---|
| mistralai/Mistral-Small-4-119B-2603 | 03/2026 | 119B total / 6.5B active (MoE) | 256K tokens | |
| mistralai/Devstral-Small-2-24B-Instruct-2512 | 12/2025 | 24B (dense) | 256K tokens |
|
| mistralai/Mistral-Medium-3.5-128B | 05/2026 | 128B (dense) | 256K tokens |
|
| mistralai/Devstral-2-123B-Instruct-2512 | 12/2025 | 123B (dense) | 256K tokens |
|
GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.
Technical classification, not legal advice. The licence text of the deployed model version is authoritative.
Open LLM licences comparedThe recommendation follows from the weight size plus headroom for context and runtime.
| Model | Format | Weights | Minimum recommended | Note |
|---|---|---|---|---|
| Mistral Small 4 | NVFP4 | 70.8 GB | Managed GPU Server 96 | |
| Mistral Small 4 | FP8 | 120.9 GB | Managed GPU Server 192 | |
| Devstral Small 2 | FP8 | 25.8 GB | Managed GPU Server 96 | |
| Mistral Medium 3.5 | FP8 | 133.6 GB | Managed GPU Server 192 | Licence note: no rights above USD 20 million consolidated monthly revenue. |
| Devstral 2 | FP8 | 128.2 GB | Managed GPU Server 192 | Licence note: no rights above USD 20 million consolidated monthly revenue. |
Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.
The FP8 sizes are the sum of the weight files of one format. The Mistral repositories also contain the weights a second time in consolidated format, which is not loaded twice in operation.
1 × RTX PRO 4000 Blackwell, 24 GB GPU memory
€699 excl. VAT / month, cancellable monthly
€499 excl. VAT one-time setup
View configurationRelevant for this family
1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory
€1,799 excl. VAT / month, cancellable monthly
€999 excl. VAT one-time setup
View configurationRelevant for this family
2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory
Price and term on request
View configuration4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory
Price and term on request
View configurationThe Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.
Architecture, quantisation and distribution across several GPUs.
Mistral Small 4, Mistral Medium 3.5 and Devstral Small 2 use Mistral3ForConditionalGeneration, an architecture supported by vLLM. Mistral recommends a current version of the mistral_common package for serving.
Mistral officially publishes Mistral Small 4 in FP8 and NVFP4. The NVFP4 version occupies around 71 GB and is therefore the variant for the Managed GPU Server 96. In the model card, Mistral uses two GPUs for both formats.
Mistral Small 4 uses multi-head latent attention. vLLM needs a suitable attention backend for it, which we verify with the deployed version on the RTX PRO 6000 Blackwell.
Mistral Medium 3.5 and Devstral 2 have 96 attention heads and 8 KV heads and can be split across two or four GPUs. The vendor examples use eight GPUs; for the FP8 weights of around 130 GB we recommend at least two RTX PRO 6000.
The Managed GPU Server is an operated model environment, not an empty server.
An agreed model in the agreed quantisation, served through a managed vLLM inference layer.
Applications and coding clients connect to the server via base URL and API key.
A managed chat interface for teams using the model without their own application.
Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.
Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.
Data processing agreement, documented configuration and a personal point of contact.
Scope, prices and multi-GPU servers are on the product page.
View Managed GPU ServerQwen hosting
Qwen hosting in Germany: Qwen3.8-27B, Qwen3.5-122B and Qwen3-Coder-Next on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familygpt-oss hosting
gpt-oss hosting in Germany: OpenAI's gpt-oss-120b and gpt-oss-20b on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyGemma hosting
Gemma hosting in Germany: Google's Gemma 4 31B, 26B A4B and 12B on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyDevstral hosting
Devstral hosting in Germany: Mistral AI's Devstral Small 2 and Devstral 2 for coding agents, with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyDeepSeek hosting
DeepSeek hosting in Germany: DeepSeek-V4-Flash on a dedicated four-GPU server with vLLM, under the MIT licence and with no data path to the DeepSeek API.
View familyLlama hosting
Llama hosting in Germany: why Llama 4 is ruled out by its licence for companies in the EU, how Llama 3.3 70B is operated and which alternatives exist.
View familyModel assessment
Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.
Licence, server tier, data path and operation
Mistral Small 4 in NVFP4 and Devstral Small 2 in FP8. Both are licensed under Apache 2.0. Mistral Small 4 in FP8, Mistral Medium 3.5 and Devstral 2 need at least the Managed GPU Server 192.
Only if your company's consolidated global monthly revenue did not exceed USD 20 million in the preceding month. Above that the Modified MIT License grants no rights; Mistral AI grants commercial licences on request. Mistral Small 4 and Devstral Small 2 are under Apache 2.0 without a threshold.
The origin of the model does not determine data protection. What matters is where and how it runs. With us, the model runs on a dedicated server in a German data centre, under a data processing agreement and without a connection to the Mistral API.
Devstral Small 2 is a dense model with 24B parameters and runs on the Managed GPU Server 96. Qwen3-Coder-Next is larger and needs at least the Managed GPU Server 192. We compare which model gives better results on your repositories using your tasks.
NVFP4 stores the weights mostly in 4 bits and occupies around 71 GB; FP8 occupies around 121 GB. For NVFP4 the minimum recommended tier is the Managed GPU Server 96, for FP8 the Managed GPU Server 192. We check quality differences using examples from your use case.
Yes. vLLM provides an OpenAI-compatible API, and Open WebUI is available as a chat interface. Coding clients with an OpenAI-compatible interface connect directly.
The entry point for Mistral is the Managed GPU Server 96, the minimum recommended tier for Mistral Small 4 and Devstral Small 2: €1,799 excl. VAT / month plus €999 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. Larger variants need a multi-GPU server, priced on request. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.