WZ-IT Logo

Running Gemma 4 with vLLM: which 4-bit format works (W4A16, not GGUF)

Timo WevelsiepTimo Wevelsiep•Updated: 03.10.2026

Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.

Want to run Gemma 4 in production with vLLM and pick the right GPU tier?

The Gemma hosting page shows which Gemma 4 checkpoints we run on which tier. On the Managed GPU Server with 24 GB VRAM from €699 excl. VAT per month, Gemma 4 12B runs as a W4A16 checkpoint; Gemma 4 31B and 26B A4B are planned for the 96 GB VRAM tier.

Explore Gemma hosting · Explore Managed GPU Server

If you want to run Gemma 4 at 4 bits, Google offers several official variants. The most visible one is the Q4_0 GGUF file, which is also the smallest. It is not intended for vLLM, though: Google provides the GGUF files for llama.cpp and the compressed-tensors checkpoints with W4A16 for vLLM (Google: QAT for Gemma 4). This article explains the formats, lists sizes and GPU tiers and covers the pitfalls at start-up. Model and file details as of October 2026.

Table of contents

The problem: the wrong 4-bit file

In June 2026, Google released checkpoints trained with quantization-aware training (QAT) for Gemma 4. QAT simulates quantisation during training, so the model loses less quality at 4 bits than with quantisation applied afterwards (Google: QAT for Gemma 4).

The QAT weights come in several packages. According to the model card, these are:

  • GGUF (Q4_0) for E2B, E4B, 12B, 26B A4B and 31B, ready for the llama.cpp ecosystem.
  • Compressed tensors (W4A16) for E2B, E4B, 12B and 31B, for native inference with vLLM.
  • Unquantised QAT checkpoints in half precision, as a starting point for your own conversions.
  • Mobile format for E2B and E4B.

Searching for "Gemma 4 4-bit" often leads to the GGUF repositories first, because they have the smallest files and are the only ready-quantised variant for 26B A4B. In a vLLM setup, that is the wrong starting point.

The formats at a glance

Format What is stored Typical runtime
BF16 16-bit weights, original precision vLLM, Transformers, SGLang
FP8 8-bit weights (and usually activations) vLLM, SGLang
W4A16 (compressed-tensors) 4-bit weights, 16-bit activations vLLM
GGUF (Q4_0) one file with 4-bit weights, tokenizer and metadata llama.cpp, Ollama, LM Studio

W4A16 means 4-bit weights (W4) and 16-bit activations (A16). The Gemma 4 checkpoints use 4-bit integer weights with a group size of 32 in the compressed-tensors format, which vLLM reads natively (vLLM recipe Gemma 4). According to the quantization_config in config.json, the embedding matrix and the vision encoder stay in 16 bits.

GGUF is the container format of llama.cpp. It is designed for single-user operation and mixed CPU/GPU inference. For Gemma 4, a separate mmproj file for image processing sits next to the language model.

The role of FP8, NVFP4, AWQ and GPTQ in general is covered in LLM quantization.

Which checkpoint on which runtime

Sizes are the sum of the weight files on Hugging Face, in GB (10⁹ bytes). They do not include KV cache.

Checkpoint Format Weights Runtime Minimum recommended at WZ-IT
gemma-4-31B-it BF16 62.5 GB vLLM Managed GPU Server 96
gemma-4-31B-it-qat-w4a16-ct W4A16 23.3 GB vLLM Managed GPU Server 96
gemma-4-31B-it-qat-q4_0-gguf GGUF Q4_0 17.7 GB + 1.2 GB mmproj llama.cpp, Ollama not intended for vLLM
gemma-4-26B-A4B-it BF16 51.6 GB vLLM Managed GPU Server 96
gemma-4-26B-A4B-it-qat-q4_0-gguf GGUF Q4_0 14.4 GB + 1.2 GB mmproj llama.cpp, Ollama not intended for vLLM
gemma-4-12B-it BF16 23.9 GB vLLM Managed GPU Server 96
gemma-4-12B-it-qat-w4a16-ct W4A16 10.3 GB vLLM Managed GPU Server 24
gemma-4-12B-it-qat-q4_0-gguf GGUF Q4_0 7.0 GB + 0.2 GB mmproj llama.cpp, Ollama not intended for vLLM

The Managed GPU Server 24 has an NVIDIA RTX PRO 4000 Blackwell with 24 GB, the Managed GPU Server 96 an RTX PRO 6000 Blackwell with 96 GB. Gemma 4 31B W4A16 is listed at 96 GB although its weights are below 24 GB: by default vLLM uses 92 percent of GPU memory (engine arguments), and after weights, activations and CUDA graphs there is hardly any room left for KV cache on 24 GB. The vLLM recipe page lists 19.8 GB for the loaded model and 8.3 GB for Gemma 4 12B W4A16 (vLLM recipe Gemma 4). How weights, cache and overhead add up is shown in Sizing GPU and VRAM for LLMs.

Starting Gemma 4 with W4A16 in vLLM

vLLM reads the quantisation from the checkpoint's configuration. No --quantization flag is needed for the W4A16 checkpoints (vLLM recipe Gemma 4). A minimal start for Gemma 4 12B on a 24 GB GPU:

vllm serve google/gemma-4-12B-it-qat-w4a16-ct \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90 \
  --limit-mm-per-prompt '{"image": 2, "audio": 0}'
  • --max-model-len limits the context per request. Without it, the maximum length from the model configuration applies (engine arguments).
  • --gpu-memory-utilization sets the share of GPU memory vLLM uses for weights and KV cache. The recipe recommends 0.85 to 0.95.
  • --limit-mm-per-prompt limits images and audio per request. Setting audio to 0 saves the memory for the audio embedder of the 12B model.

For function calling and thinking mode, the recipe adds --enable-auto-tool-choice, --tool-call-parser gemma4 and --reasoning-parser gemma4. Gemma 4 31B W4A16 starts with the same arguments, then on a 96 GB GPU. The API is OpenAI-compatible; Open WebUI or other clients call the /v1/chat/completions endpoint. How vLLM works in general is described in What is vLLM?.

Why GGUF is not the standard route in vLLM

vLLM can load GGUF files, but not on a par with its own formats. The GGUF page of the vLLM documentation states three points:

  • The support is highly experimental and under-optimised and may be incompatible with other features.
  • It has moved to the separate vllm-gguf-plugin, which must be installed in addition.
  • The tokenizer should come from the base model, because converting it from GGUF is slow and unstable, especially with large vocabularies. Gemma 4 has a vocabulary of 262,144 tokens.

The background is an RFC in the vLLM project to move GGUF and bitsandbytes out of the core. It estimates GGUF's share of usage at around 0.1 percent and notes that GGUF models often run faster with llama.cpp. The plugin does list Gemma 4 among its tested models, but with a Q4_K_M quantisation and a BF16 projector, not with Google's QAT Q4_0 file.

For a production vLLM server with several users, this means: the W4A16 checkpoint uses vLLM's optimised kernels and regular loading path. The GGUF file belongs with llama.cpp, Ollama or LM Studio. The difference between the two engines is covered in vLLM vs. Ollama.

Special case Gemma 4 26B A4B

There is no W4A16 checkpoint for the mixture-of-experts model 26B A4B. According to the vLLM recipe, the model loses too much quality at 4 bits because of its small expert dimension (128 experts with an intermediate size of 704). For vLLM, two routes remain:

  • BF16 with around 51.6 GB of weights, on a 96 GB GPU.
  • Online quantisation at load time with --quantization int8_per_channel_weight_only. The recipe states around 47 percent memory savings with negligible quality loss; no checkpoint is required (vLLM: online quantization).

At around 27 GB of weights, the INT8 variant does not fit on 24 GB either. For 26B A4B, the 96 GB tier remains the starting point. Only around 3.8 billion parameters are active per token, so the model responds faster than 31B.

Pitfalls: context length and image inputs

Context length. Gemma 4 12B, 26B A4B and 31B support up to 256K tokens; the config.json of 31B specifies 262,144 tokens. Without --max-model-len, vLLM adopts this value and has to provide cache for at least one sequence of that length. On tight GPUs, start-up then fails. A fixed value such as 32768 works for chat and document questions. Alternatively, --max-model-len auto picks the largest length that fits in memory (engine arguments).

Gemma 4 relieves the KV cache through hybrid attention. For 31B, according to config.json, only 10 of 60 layers use global attention, the other 50 a sliding window of 1,024 tokens. The cache therefore grows more slowly than in a model with full attention in every layer, but it still grows with context and the number of concurrent sessions. An FP8 KV cache with --kv-cache-dtype fp8 halves it further; check the accuracy before going into production.

Image inputs. All three models accept text and images, 12B also audio. Depending on the budget, each image occupies 70, 140, 280, 560 or 1,120 tokens of context; the default is 280, configurable with --mm-processor-kwargs '{"max_soft_tokens": 560}' (vLLM recipe Gemma 4). Higher budgets help with OCR and small print but consume context accordingly. Without a limit, vLLM allows up to 999 inputs per modality and request. If you only process text, set --language-model-only; vLLM then skips the multimodal memory reservation.

In the GGUF variant, image capabilities live in the separate mmproj file. Loading only the language model file gives you no image understanding.

Licence: Apache 2.0 from Gemma 4 onwards

Gemma 4 is licensed under the Apache License 2.0, including the QAT checkpoints. Gemma 3 and older versions remain under the Gemma Terms of Use with their own Prohibited Use Policy. When moving from Gemma 3 to Gemma 4, update your licence review accordingly. A comparison of open model licences is available in Open LLM licences for commercial use.

Our approach at WZ-IT

On our Managed GPU Servers we run Gemma 4 with vLLM. We use the checkpoint Google intends for vLLM and set context length, number of concurrent sessions and image budget according to your use case:

  1. Choose model and format: Gemma 4 12B W4A16 on 24 GB, Gemma 4 31B (BF16 or W4A16) or 26B A4B on 96 GB.
  2. Check memory: calculate weights, KV cache at the required context length and overhead against the GPU; the available cache is shown in the vLLM start-up log.
  3. Test with your own examples: answers in your language, documents and images from your daily work, before the model goes into production.

Which Gemma 4 models we use on which tier is shown on the Gemma hosting page. The LLM hosting page compares the operating options.

Further reading: Which LLM to self-host? places Gemma alongside other open models, Multi-GPU inference without NVLink describes the step to several GPUs.

Sources

Rather have it operated?

You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.

Enquiry

Assess local AI for your use case

Start with the AI Cube or have us assess a custom AI platform, knowledge connection, or integration.

How should we get back to you?

Frequently Asked Questions

Answers to the most important questions

The QAT checkpoints in compressed-tensors format with the suffix -qat-w4a16-ct, such as google/gemma-4-31B-it-qat-w4a16-ct and google/gemma-4-12B-it-qat-w4a16-ct. Google describes them as the format for native inference with vLLM. vLLM detects the quantisation from config.json, so no --quantization flag is needed.

No. GGUF is the format of llama.cpp and tools built on it such as Ollama and LM Studio. vLLM now loads GGUF only through the separate vllm-gguf-plugin, the vLLM documentation describes the support as highly experimental and under-optimised, and it may be incompatible with other features. For production use with vLLM, the W4A16 checkpoint is the intended route.

The weight files of Gemma 4 31B W4A16 total around 23.3 GB, those of Gemma 4 12B W4A16 around 10.3 GB. For comparison, Gemma 4 31B in BF16 is around 62.5 GB. KV cache, activations and runtime overhead come on top.

Not in a useful way for operation. The weights occupy almost all of the memory, leaving hardly any room for context and several users. We recommend at least the Managed GPU Server 96 for it. On 24 GB, Gemma 4 12B W4A16 is the right choice.

No. Google publishes W4A16 for E2B, E4B, 12B and 31B. According to the vLLM recipe, the MoE model loses too much quality at 4 bits because of its small expert dimension. For 26B A4B, vLLM instead lists online quantisation with int8_per_channel_weight_only or running it in BF16 at around 51.6 GB.

Usually because of the context length. Without --max-model-len, vLLM takes the 262,144 tokens from the model configuration and needs KV cache for at least one full sequence. A fixed value such as 32768, or the auto setting that picks the largest length that fits, resolves this.

Yes. Depending on the budget, each image occupies 70 to 1,120 tokens of context, with 280 as the default. vLLM also reserves memory for multimodal processing at start-up. If you only need text, disable image and audio inputs with --language-model-only or --limit-mm-per-prompt.

Yes. Gemma 4 is licensed under the Apache License 2.0, without a revenue or user threshold. Gemma 3 and older versions remain under the Gemma Terms of Use with their own Prohibited Use Policy.

More on Local AI for Business

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.