You want to use DeepSeek without sending data to the DeepSeek API in China. We run the open weights on a dedicated multi-GPU server in a German data centre.
Companies worldwide trust WZ-IT
DeepSeek publishes its models as open weights under the MIT licence. The current V4 generation is designed for long contexts of up to 1 million tokens and agentic tasks. The models are large: DeepSeek-V4-Flash alone has 284B parameters.
The key distinction is between weights and the vendor API. Using the DeepSeek app or DeepSeek API sends data to the vendor's servers. With hosting, the downloaded weights run on your server in Germany, with no connection to DeepSeek.
The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.
A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.
Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.
Vendor specifications and weight file size per format, with source.
| Checkpoint | Released | Parameters | Context (vendor) | Weights per format |
|---|---|---|---|---|
| deepseek-ai/DeepSeek-V4-Flash-0731 | 07/2026 | 284B total / 13B active (MoE) | 1M tokens |
|
| deepseek-ai/DeepSeek-V4.1-Flash | 09/2026 | 552B backbone / 8B or 16B active | 1M tokens |
GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.
Technical classification, not legal advice. The licence text of the deployed model version is authoritative.
Open LLM licences comparedThe recommendation follows from the weight size plus headroom for context and runtime.
| Model | Format | Weights | Minimum recommended | Note |
|---|---|---|---|---|
| DeepSeek-V4-Flash-0731 | FP4/FP8 | 166.9 GB | Managed GPU Server 384 | On 192 GB only around 25 GB would remain for context and runtime after the weights. |
| DeepSeek-V4.1-Flash | FP4/FP8 | 510.3 GB | None of our tiers | The weights exceed our largest tier of 384 GB. We do not offer this model. |
Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.
1 × RTX PRO 4000 Blackwell, 24 GB GPU memory
€699 excl. VAT / month, cancellable monthly
€499 excl. VAT one-time setup
View configuration1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory
€1,799 excl. VAT / month, cancellable monthly
€999 excl. VAT one-time setup
View configuration2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory
Price and term on request
View configurationRelevant for this family
4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory
Price and term on request
View configurationThe Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.
Architecture, quantisation and distribution across several GPUs.
DeepSeek-V4-Flash uses DeepseekV4ForCausalLM, an architecture supported by vLLM. DeepSeek-V4.1-Flash uses a new encoder-decoder architecture that was not on vLLM's list of supported models at the time of checking.
The expert weights are stored in 4 bits, other parts in FP8. The vendor examples refer to GB300 systems. We verify which kernels run on the RTX PRO 6000 Blackwell with the deployed vLLM version before the proposal.
With 64 attention heads, V4-Flash can be split across two or four GPUs using tensor parallelism, or using expert parallelism. For V4-Flash we recommend at least four GPUs.
DeepSeek specifies a context of 1 million tokens. Every long request occupies memory for the KV cache. We set a sensible context length on the server based on your requests. 1 million tokens are not automatically part of the configuration.
The Managed GPU Server is an operated model environment, not an empty server.
An agreed model in the agreed quantisation, served through a managed vLLM inference layer.
Applications and coding clients connect to the server via base URL and API key.
A managed chat interface for teams using the model without their own application.
Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.
Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.
Data processing agreement, documented configuration and a personal point of contact.
Scope, prices and multi-GPU servers are on the product page.
View Managed GPU ServerQwen hosting
Qwen hosting in Germany: Qwen3.8-27B, Qwen3.5-122B and Qwen3-Coder-Next on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familygpt-oss hosting
gpt-oss hosting in Germany: OpenAI's gpt-oss-120b and gpt-oss-20b on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyGemma hosting
Gemma hosting in Germany: Google's Gemma 4 31B, 26B A4B and 12B on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyMistral hosting
Mistral hosting in Germany: Mistral Small 4, Devstral Small 2 and Mistral Medium 3.5 on a dedicated GPU, with vLLM, licence review and operation under a DPA.
View familyDevstral hosting
Devstral hosting in Germany: Mistral AI's Devstral Small 2 and Devstral 2 for coding agents, with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyLlama hosting
Llama hosting in Germany: why Llama 4 is ruled out by its licence for companies in the EU, how Llama 3.3 70B is operated and which alternatives exist.
View familyModel assessment
Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.
Licence, server tier, data path and operation
Not with hosting. Data only goes to DeepSeek if you use the DeepSeek app or the DeepSeek API. We run the downloaded weights on a dedicated server in a German data centre. There the model has no connection to the vendor, and requests and answers stay on the server.
Operation can be set up in compliance with data protection rules: the weights run on a dedicated server in a German data centre, WZ-IT signs a data processing agreement, and no data is transferred to DeepSeek or to a third country. The GDPR obligations for your own processing, such as the legal basis and the record of processing activities, remain with you. The DeepSeek app and API are different: there, data is transferred to the provider in China.
The minimum recommended tier is the Managed GPU Server 384 with four RTX PRO 6000 Blackwell Max-Q. The weights occupy around 167 GB; with two GPUs too little memory would remain for context and concurrent requests.
No. The V4.1-Flash weights occupy more than 500 GB and therefore exceed our largest tier. In addition, the architecture was not on vLLM's list of supported models at the time of checking.
Yes. The repository and model weights are licensed under MIT, without a revenue or user threshold. The licence notice must be included when redistributing.
The current V4 generation has no model that can run on a single 96 GB GPU. For one GPU we recommend models such as Qwen3.8-27B or gpt-oss-120b, which we cover on the respective family pages.
Model weights in safetensors format contain only numbers, no executable code. vLLM provides the inference code. We only use additional scripts from the repository after review, and the server's outbound connections are agreed.
The minimum recommended tier for DeepSeek-V4-Flash-0731 is the Managed GPU Server 384. Multi-GPU server pricing is available on request, the term is set out in the proposal. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.