WZ-IT Logo
Gemma hosting

Gemma Hosting in Germany: Run Gemma 4 on a Dedicated GPU

You want a multilingual model with image input under Apache 2.0. We run Gemma 4 on a dedicated GPU server in a German data centre.

Operated in a German data centreOperation under a DPAvLLM and OpenAI-compatible API

Companies worldwide trust WZ-IT

Reviews

Companies worldwide trust WZ-IT

Stadtwerke BrühlDGHO e.V.ABCO Water SystemsGolem.deEVADXBnextGYMAInergyml&sOdiseo SolutionsAnnotaARGESweetConnect GmbH
Starting point

Why run Gemma yourself

Gemma is Google DeepMind's open-weight family. Gemma 4 is available as dense models (31B, 12B) and as a mixture-of-experts model (26B A4B), with image input, more than 140 languages and up to 256K context. With Gemma 4, Google moved the licence to Apache 2.0.

Many companies find Gemma 4 attractive because Google publishes official 4-bit checkpoints trained with quantization-aware training. This makes smaller GPU tiers usable without relying on community quantisations.

  • Data stays in Germany

    The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.

  • Predictable costs

    A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.

  • Model version under control

    Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.

Model facts

Gemma checkpoints at a glance

Vendor specifications and weight file size per format, with source.

CheckpointReleasedParametersContext (vendor)Weights per format
google/gemma-4-31B-it04/202630.7B (dense)256K tokens
google/gemma-4-26B-A4B-it04/202625.2B total / 3.8B active (MoE)256K tokens
google/gemma-4-12B-it05/202611.95B (dense)256K tokens

GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.

Licence

Gemma: licence and commercial use

Licence
Apache License 2.0 (Gemma 4) (Licence text)
Commercial use
Yes, without a revenue or user threshold.
Obligations and thresholds
The Gemma 4 licence page contains the Apache 2.0 text without additional terms of use. Gemma 3 and older versions remain under the Gemma Terms of Use with the Prohibited Use Policy.
Checked
As of October 2026

Technical classification, not legal advice. The licence text of the deployed model version is authoritative.

Open LLM licences compared
Server tier

Minimum recommended server tier for Gemma

The recommendation follows from the weight size plus headroom for context and runtime.

ModelFormatWeightsMinimum recommendedNote
Gemma 4 31BBF1662.5 GBManaged GPU Server 96
Gemma 4 31BQAT W4A1623.3 GBManaged GPU Server 96The weights occupy almost all of the 24 GB. Context and runtime require the next tier.
Gemma 4 26B A4BBF1651.6 GBManaged GPU Server 96
Gemma 4 12BBF1623.9 GBManaged GPU Server 96
Gemma 4 12BQAT W4A1610.3 GBManaged GPU Server 24

Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

The QAT Q4_0 files in GGUF format are intended for llama.cpp and tools built on it. Our server tiers use vLLM, so we list the W4A16 checkpoints in compressed-tensors format there. At the time of checking, Google does not publish a W4A16 checkpoint for Gemma 4 26B A4B.

The Managed GPU Server tiers

Relevant for this family

Managed GPU Server 24

1 × RTX PRO 4000 Blackwell, 24 GB GPU memory

€699 excl. VAT / month, cancellable monthly

€499 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 96

1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory

€1,799 excl. VAT / month, cancellable monthly

€999 excl. VAT one-time setup

View configuration

Managed GPU Server 192

2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory

Price and term on request

View configuration

Managed GPU Server 384

4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory

Price and term on request

View configuration

The Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.

vLLM

Running Gemma with vLLM

Architecture, quantisation and distribution across several GPUs.

Architectures

Gemma 4 31B and 26B A4B use Gemma4ForConditionalGeneration, Gemma 4 12B the encoder-free variant Gemma4UnifiedForConditionalGeneration. Both are listed among the models supported by vLLM.

Quantisation

For vLLM, Google publishes the QAT checkpoints in compressed-tensors format (W4A16: 4-bit weights, 16-bit activations). The Q4_0 GGUF files target llama.cpp.

Hybrid attention

Gemma 4 alternates between local sliding-window attention and global attention. This reduces the memory needed for long contexts but does not replace sizing of context length and parallelism.

Tensor parallelism

Gemma 4 31B has 32 attention heads, 26B A4B and 12B have 16. One GPU is sufficient for the listed checkpoints. Two GPUs are an option for long contexts or several models on one server.

Operated by WZ-IT

What operation includes

The Managed GPU Server is an operated model environment, not an empty server.

Model and vLLM

An agreed model in the agreed quantisation, served through a managed vLLM inference layer.

OpenAI-compatible API

Applications and coding clients connect to the server via base URL and API key.

Open WebUI

A managed chat interface for teams using the model without their own application.

24/7 monitoring

Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.

Updates with CVE assessment

Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.

DPA and documentation

Data processing agreement, documented configuration and a personal point of contact.

Scope, prices and multi-GPU servers are on the product page.

View Managed GPU Server

Model assessment

Have your model and server tier assessed

Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.

Which server tier should be assessed?

The selection is non-binding. We confirm model fit, licence and term before a proposal.

How should we get back to you?

We usually respond within one business day. Please do not send access credentials yet.

Frequently asked questions about Gemma hosting

Licence, server tier, data path and operation

Gemma 4 12B as a QAT W4A16 checkpoint is the variant for the Managed GPU Server 24. For higher quality, Gemma 4 31B in BF16 or W4A16 is the next step, with the Managed GPU Server 96 as the minimum recommended tier.

The Gemma 4 licence page contains the text of the Apache License 2.0 without reference to additional terms of use. The Gemma Terms of Use with the Prohibited Use Policy cover Gemma 3 and older versions.

Gemma 4 12B with W4A16 does; the weights occupy around 10 GB. Gemma 4 31B with W4A16 occupies around 23 GB and leaves no room for context on 24 GB, so we recommend at least the Managed GPU Server 96. The smaller GGUF files are intended for llama.cpp, not vLLM.

For the weights alone: Gemma 4 12B needs around 10 GB with W4A16 and around 24 GB in BF16, Gemma 4 26B A4B around 52 GB in BF16, Gemma 4 31B around 23 GB with W4A16 and around 63 GB in BF16. Context and concurrent requests come on top. The minimum recommended tier is therefore the Managed GPU Server 24 for Gemma 4 12B with W4A16 and the Managed GPU Server 96 for all other variants.

No. The weights run on your dedicated server in a German data centre, with no connection to Google services. This applies to the hosted model, not to tools you additionally connect to external services.

Yes. All Gemma 4 models listed here accept text and images and output text. Image inputs take up additional context, which we account for in sizing.

Google states support for more than 140 languages. We test answers with your own examples before the model goes into production.

The entry point for Gemma is the Managed GPU Server 24, the minimum recommended tier for Gemma 4 12B: €699 excl. VAT / month plus €499 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.