WZ-IT Logo
DeepSeek hosting

DeepSeek Hosting in Germany: DeepSeek-V4-Flash on Dedicated GPUs

You want to use DeepSeek without sending data to the DeepSeek API in China. We run the open weights on a dedicated multi-GPU server in a German data centre.

Operated in a German data centreOperation under a DPAvLLM and OpenAI-compatible API

Companies worldwide trust WZ-IT

Reviews

Companies worldwide trust WZ-IT

Stadtwerke BrühlDGHO e.V.ABCO Water SystemsGolem.deEVADXBnextGYMAInergyml&sOdiseo SolutionsAnnotaARGESweetConnect GmbH
Starting point

Why run DeepSeek yourself

DeepSeek publishes its models as open weights under the MIT licence. The current V4 generation is designed for long contexts of up to 1 million tokens and agentic tasks. The models are large: DeepSeek-V4-Flash alone has 284B parameters.

The key distinction is between weights and the vendor API. Using the DeepSeek app or DeepSeek API sends data to the vendor's servers. With hosting, the downloaded weights run on your server in Germany, with no connection to DeepSeek.

  • Data stays in Germany

    The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.

  • Predictable costs

    A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.

  • Model version under control

    Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.

Model facts

DeepSeek checkpoints at a glance

Vendor specifications and weight file size per format, with source.

CheckpointReleasedParametersContext (vendor)Weights per format
deepseek-ai/DeepSeek-V4-Flash-073107/2026284B total / 13B active (MoE)1M tokens
deepseek-ai/DeepSeek-V4.1-Flash09/2026552B backbone / 8B or 16B active1M tokens

GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.

Licence

DeepSeek: licence and commercial use

Licence
MIT License (Licence text)
Commercial use
Yes, without a revenue or user threshold. The MIT licence covers the repository and the model weights.
Obligations and thresholds
Include the copyright and licence notice when redistributing. No further terms of use in the repository.
Checked
As of October 2026

Technical classification, not legal advice. The licence text of the deployed model version is authoritative.

Open LLM licences compared
Server tier

Minimum recommended server tier for DeepSeek

The recommendation follows from the weight size plus headroom for context and runtime.

ModelFormatWeightsMinimum recommendedNote
DeepSeek-V4-Flash-0731FP4/FP8166.9 GBManaged GPU Server 384On 192 GB only around 25 GB would remain for context and runtime after the weights.
DeepSeek-V4.1-FlashFP4/FP8510.3 GBNone of our tiersThe weights exceed our largest tier of 384 GB. We do not offer this model.

Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

The Managed GPU Server tiers

Managed GPU Server 24

1 × RTX PRO 4000 Blackwell, 24 GB GPU memory

€699 excl. VAT / month, cancellable monthly

€499 excl. VAT one-time setup

View configuration

Managed GPU Server 96

1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory

€1,799 excl. VAT / month, cancellable monthly

€999 excl. VAT one-time setup

View configuration

Managed GPU Server 192

2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory

Price and term on request

View configuration

Relevant for this family

Managed GPU Server 384

4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory

Price and term on request

View configuration

The Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.

vLLM

Running DeepSeek with vLLM

Architecture, quantisation and distribution across several GPUs.

Architecture

DeepSeek-V4-Flash uses DeepseekV4ForCausalLM, an architecture supported by vLLM. DeepSeek-V4.1-Flash uses a new encoder-decoder architecture that was not on vLLM's list of supported models at the time of checking.

Weight format

The expert weights are stored in 4 bits, other parts in FP8. The vendor examples refer to GB300 systems. We verify which kernels run on the RTX PRO 6000 Blackwell with the deployed vLLM version before the proposal.

Distribution

With 64 attention heads, V4-Flash can be split across two or four GPUs using tensor parallelism, or using expert parallelism. For V4-Flash we recommend at least four GPUs.

Context

DeepSeek specifies a context of 1 million tokens. Every long request occupies memory for the KV cache. We set a sensible context length on the server based on your requests. 1 million tokens are not automatically part of the configuration.

Operated by WZ-IT

What operation includes

The Managed GPU Server is an operated model environment, not an empty server.

Model and vLLM

An agreed model in the agreed quantisation, served through a managed vLLM inference layer.

OpenAI-compatible API

Applications and coding clients connect to the server via base URL and API key.

Open WebUI

A managed chat interface for teams using the model without their own application.

24/7 monitoring

Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.

Updates with CVE assessment

Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.

DPA and documentation

Data processing agreement, documented configuration and a personal point of contact.

Scope, prices and multi-GPU servers are on the product page.

View Managed GPU Server

Model assessment

Have your model and server tier assessed

Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.

Which server tier should be assessed?

The selection is non-binding. We confirm model fit, licence and term before a proposal.

How should we get back to you?

We usually respond within one business day. Please do not send access credentials yet.

Frequently asked questions about DeepSeek hosting

Licence, server tier, data path and operation

Not with hosting. Data only goes to DeepSeek if you use the DeepSeek app or the DeepSeek API. We run the downloaded weights on a dedicated server in a German data centre. There the model has no connection to the vendor, and requests and answers stay on the server.

Operation can be set up in compliance with data protection rules: the weights run on a dedicated server in a German data centre, WZ-IT signs a data processing agreement, and no data is transferred to DeepSeek or to a third country. The GDPR obligations for your own processing, such as the legal basis and the record of processing activities, remain with you. The DeepSeek app and API are different: there, data is transferred to the provider in China.

The minimum recommended tier is the Managed GPU Server 384 with four RTX PRO 6000 Blackwell Max-Q. The weights occupy around 167 GB; with two GPUs too little memory would remain for context and concurrent requests.

No. The V4.1-Flash weights occupy more than 500 GB and therefore exceed our largest tier. In addition, the architecture was not on vLLM's list of supported models at the time of checking.

Yes. The repository and model weights are licensed under MIT, without a revenue or user threshold. The licence notice must be included when redistributing.

The current V4 generation has no model that can run on a single 96 GB GPU. For one GPU we recommend models such as Qwen3.8-27B or gpt-oss-120b, which we cover on the respective family pages.

Model weights in safetensors format contain only numbers, no executable code. vLLM provides the inference code. We only use additional scripts from the repository after review, and the server's outbound connections are agreed.

The minimum recommended tier for DeepSeek-V4-Flash-0731 is the Managed GPU Server 384. Multi-GPU server pricing is available on request, the term is set out in the proposal. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.