WZ-IT Logo
Devstral hosting

Devstral Hosting in Germany: Run Mistral's Coding Models Yourself

You want to run coding agents and IDE assistants on your own model without sending source code to a third-party API. We run Devstral on a dedicated GPU server in a German data centre.

Operated in a German data centreOperation under a DPAvLLM and OpenAI-compatible API

Companies worldwide trust WZ-IT

Reviews

Companies worldwide trust WZ-IT

Stadtwerke BrühlDGHO e.V.ABCO Water SystemsGolem.deEVADXBnextGYMAInergyml&sOdiseo SolutionsAnnotaARGESweetConnect GmbH
Starting point

Why run Devstral yourself

Devstral is Mistral AI's coding line. The models are designed for software development agents: exploring a codebase, editing several files and calling tools. Devstral Small 2 has 24B parameters, also accepts images and is licensed under Apache 2.0. Devstral 2 has 123B parameters and is licensed under a Modified MIT License with a revenue threshold.

Coding agents send large parts of a codebase to the model as context. On your own server, source code, diffs and test output stay in the German data centre. IDE extensions and agent clients connect to the server through vLLM's OpenAI-compatible API. We cover the general-purpose models Mistral Small 4 and Mistral Medium 3.5 on the Mistral hosting page.

  • Data stays in Germany

    The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.

  • Predictable costs

    A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.

  • Model version under control

    Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.

Model facts

Devstral checkpoints at a glance

Vendor specifications and weight file size per format, with source.

CheckpointReleasedParametersContext (vendor)Weights per format
mistralai/Devstral-Small-2-24B-Instruct-251212/202524B (dense, with vision)256K tokens
mistralai/Devstral-2-123B-Instruct-251212/2025123B (dense, text only)256K tokens

GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.

Licence

Devstral: licence and commercial use

Licence
Apache 2.0 (Devstral Small 2) / Modified MIT (Devstral 2) (Licence text, Devstral Small 2 model card)
Commercial use
Devstral Small 2: yes, without a revenue or user threshold. Devstral 2: only if the consolidated global monthly revenue of your company (or your employer) did not exceed USD 20 million in the preceding month.
Obligations and thresholds
Modified MIT: include attribution and the licence notice. Above the revenue threshold the licence grants no rights, including for derived or combined models; Mistral AI grants commercial licences on request. Apache 2.0: include the licence text and notices when redistributing.
Checked
As of October 2026

Technical classification, not legal advice. The licence text of the deployed model version is authoritative.

Open LLM licences compared
Server tier

Minimum recommended server tier for Devstral

The recommendation follows from the weight size plus headroom for context and runtime.

ModelFormatWeightsMinimum recommendedNote
Devstral Small 2FP825.8 GBManaged GPU Server 96The weights alone exceed 24 GB. The headroom on 96 GB goes into context and concurrent agent runs.
Devstral 2FP8128.2 GBManaged GPU Server 192Licence note: no rights above USD 20 million consolidated monthly company revenue.

Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

Mistral publishes both models as FP8 checkpoints. The repositories also contain the weights a second time in consolidated format, which is not loaded twice in operation.

The Managed GPU Server tiers

Managed GPU Server 24

1 × RTX PRO 4000 Blackwell, 24 GB GPU memory

€699 excl. VAT / month, cancellable monthly

€499 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 96

1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory

€1,799 excl. VAT / month, cancellable monthly

€999 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 192

2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory

Price and term on request

View configuration

Managed GPU Server 384

4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory

Price and term on request

View configuration

The Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.

vLLM

Running Devstral with vLLM

Architecture, quantisation and distribution across several GPUs.

Architectures

Devstral Small 2 uses Mistral3ForConditionalGeneration with an image encoder, Devstral 2 the text-only architecture Ministral3ForCausalLM. Both are listed among the models supported by vLLM. Mistral requires mistral_common version 1.8.6 or later for serving.

Tool calling for agents

Coding agents call tools via function calling. vLLM enables this for Devstral with the mistral tool-call parser and automatic tool choice, as described in the model card. Clients with an OpenAI-compatible interface use it via base URL and API key.

Context and concurrent agents

Every agent run keeps its context in the KV cache. Long contexts and many concurrent runs share the same GPU memory. Mistral's own example serves Devstral Small 2 with the full 262,144 tokens on two GPUs. On one GPU we set context length and parallel runs to suit your team.

Tensor parallelism

Devstral Small 2 has 32 attention heads and Devstral 2 has 96, both with 8 KV heads. Both can be split across two or four GPUs. Mistral's example for Devstral 2 uses eight GPUs; for the FP8 weights of around 128 GB we recommend at least two RTX PRO 6000.

Operated by WZ-IT

What operation includes

The Managed GPU Server is an operated model environment, not an empty server.

Model and vLLM

An agreed model in the agreed quantisation, served through a managed vLLM inference layer.

OpenAI-compatible API

Applications and coding clients connect to the server via base URL and API key.

Open WebUI

A managed chat interface for teams using the model without their own application.

24/7 monitoring

Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.

Updates with CVE assessment

Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.

DPA and documentation

Data processing agreement, documented configuration and a personal point of contact.

Scope, prices and multi-GPU servers are on the product page.

View Managed GPU Server

Model assessment

Have your model and server tier assessed

Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.

Which server tier should be assessed?

The selection is non-binding. We confirm model fit, licence and term before a proposal.

How should we get back to you?

We usually respond within one business day. Please do not send access credentials yet.

Frequently asked questions about Devstral hosting

Licence, server tier, data path and operation

Devstral Small 2 is the starting point for IDE assistants and coding agents; the minimum recommended tier is the Managed GPU Server 96. Devstral 2 is the larger model; the minimum recommended tier is the Managed GPU Server 192, and its licence has a revenue threshold. Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

Only if the consolidated global monthly revenue of your company or your employer did not exceed USD 20 million in the preceding month. Above that the Modified MIT License grants no rights, including for derived models; Mistral AI grants commercial licences on request. Devstral Small 2 is under Apache 2.0 without a threshold.

Clients and IDE extensions with an OpenAI-compatible interface connect via base URL, API key and model name. In the model card, Mistral lists Mistral Vibe, Cline, Kilo Code, OpenHands and SWE-agent, among others. vLLM provides tool calling through its parser for Mistral models.

The model runs on a dedicated server in a German data centre, without a connection to the Mistral API. Many coding clients also have their own settings for providers and telemetry. We agree with you that the clients only address your own endpoint.

This depends less on the number of people than on concurrent agent runs and their context length. An agent with a large repository context occupies far more KV cache than a short chat question. We set context length and parallelism using your typical tasks and state the values in the proposal.

Devstral Small 2 is a dense model with 24B parameters and fits on one 96 GB GPU. Qwen3-Coder-Next is a larger mixture-of-experts model and needs at least the Managed GPU Server 192. We compare which model gives better results on your repositories using your tasks.

The models support it, but every long context occupies GPU memory. The full context is not automatically part of every tier. We set the context length to suit the server tier and concurrent agent runs.

The entry point for Devstral is the Managed GPU Server 96, the minimum recommended tier for Devstral Small 2: €1,799 excl. VAT / month plus €999 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. Larger variants need a multi-GPU server, priced on request. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.