LLM quantization explained: FP8, NVFP4, AWQ, GPTQ and GGUF
Timo Wevelsiep•Updated: 01.10.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Which model in which format fits your hardware? WZ-IT sets up local language models on your own infrastructure, chooses model and quantization to match GPU, context length and number of users, and checks quality with your own tasks. Explore managed GPU servers · Explore the AI Cube · Book a meeting
Whether a language model fits on a GPU and how many requests it can serve at once depends to a large extent on numerical precision. A model with 32 billion parameters takes around 66 GB in BF16 and less than a third of that as a 4-bit variant. In between lie formats with different assumptions about hardware and software: FP8, NVFP4, MXFP4, AWQ, GPTQ and GGUF. This article classifies the formats, shows which GPU accelerates which format natively and explains how the KV cache enters the calculation. As of October 2026.
Table of contents
- What quantization means for LLMs
- The formats at a glance
- FP8: the default from Ada and Hopper onwards
- NVFP4 and MXFP4: 4-bit floating point
- AWQ and GPTQ: 4-bit weights with calibration
- GGUF: the format of llama.cpp and Ollama
- Memory footprint by example
- Quantizing the KV cache
- Which format for which hardware
- Measuring quality loss instead of guessing
What quantization means for LLMs
Language models are now mostly published in BF16, that is with 16 bits per parameter. Quantization maps these values onto a format with fewer bits, typically 8 or 4. So that the small numbers cover the original value range, all methods additionally store scale factors, per tensor, per channel or per block of a few values.
Three things can be quantized, and the shorthand shows which is meant:
| Notation | Weights | Activations | Typical example |
|---|---|---|---|
| W8A8 | 8 bits | 8 bits | FP8 on Ada, Hopper, Blackwell |
| W4A16 | 4 bits | 16 bits | AWQ, GPTQ, INT4 |
| W4A4 | 4 bits | 4 bits | NVFP4 on Blackwell |
| KV cache | - | - | FP8 KV cache in vLLM, q8_0 in Ollama |
The benefit has two sides. First, the memory needed for the weights falls, so the model fits on smaller or fewer GPUs. Second, the GPU reads all active weights from memory for every generated token. Smaller data types consume less memory bandwidth, which according to the vLLM documentation on LLM Compressor often means higher throughput, especially for memory-bound workloads. Whether the computation itself also runs at low precision (W8A8, W4A4) depends on whether the GPU has Tensor Cores for it.
How weights, KV cache and overhead add up to the total VRAM requirement is worked through in Sizing GPU and VRAM for LLMs.
The formats at a glance
| Format | Bits per weight | Type | Calibration data | Mainly used with | Native compute |
|---|---|---|---|---|---|
| BF16 / FP16 | 16 | Floating point | no | all engines | all current GPUs |
| FP8 (E4M3) | 8 | Floating point | optional | vLLM | from Ada Lovelace (CC 8.9), Hopper, Blackwell |
| INT8 W8A8 | 8 | Integer | yes | vLLM | from Turing according to vLLM |
| NVFP4 | about 4.5 incl. scales | Floating point, block of 16 | yes | vLLM | Blackwell |
| MXFP4 | 4 plus scale per 32 values | Floating point, block of 32 | depends on model | gpt-oss in vLLM, llama.cpp | Blackwell |
| AWQ | about 4 plus scales | Integer, group-wise | yes | vLLM | W4A16, compute in 16 bits |
| GPTQ | 4 or 8 plus scales | Integer, group-wise | yes | vLLM | W4A16, compute in 16 bits |
| GGUF Q8_0 / Q4_K_M | 8.5 / 4.89 | Integer, block-wise | optional (imatrix) | llama.cpp, Ollama | CPU, GPU, Apple Silicon |
Sources: vLLM, FP8 W8A8, vLLM, Quantization, NVIDIA, Introducing NVFP4, llama.cpp, quantize. The GGUF values apply to Llama 3.1 8B as measured by llama.cpp; other models differ slightly.
Note: the overview table in the vLLM documentation ends with the Hopper column (as of October 2026). Blackwell is listed on the pages of the individual methods, for example FP8 W8A8 and INT4 W4A16.
FP8: the default from Ada and Hopper onwards
FP8 halves the memory requirement compared with BF16. For inference, E4M3 is the most common variant: 1 sign bit, 4 exponent bits, 3 mantissa bits, values up to ±448. The E5M2 variant has more range (up to ±57,344) but less precision (vLLM, FP8 W8A8).
| Property | Value according to the vLLM documentation |
|---|---|
| Computation in FP8 (W8A8) | GPUs with compute capability 8.9 or higher: Ada Lovelace, Hopper, Blackwell |
| Older GPUs | from compute capability 7.5 (Turing) weights only in FP8 (W8A16, Marlin) |
| Memory saving | factor 2 compared with 16 bits |
| Throughput | up to factor 1.6 |
| Accuracy | "minimal impact on accuracy" |
FP8 is the obvious starting point for vLLM on current NVIDIA GPUs, because many vendors publish official FP8 checkpoints. One example is Qwen3-8B-FP8 with block-wise FP8 quantization in blocks of 128 × 128. vLLM recognises such checkpoints from their quantization_config and loads them without further parameters. With --quantization fp8_per_tensor or fp8_per_block, vLLM also converts a BF16 model to FP8 at load time, without a pre-quantized checkpoint (vLLM, Online Quantization). For production systems a tested checkpoint, for example from llm-compressor, is the more traceable option.
NVFP4 and MXFP4: 4-bit floating point
NVFP4 is NVIDIA's 4-bit format for Blackwell. Each value is a 4-bit floating-point number (E2M1, range roughly -6 to +6). Every 16 values share a scale factor in FP8 (E4M3), and there is an additional FP32 factor per tensor. Including the scales, this comes to about 4.5 bits per value (NVIDIA, Introducing NVFP4).
| Property | NVFP4 | MXFP4 (OCP Microscaling) |
|---|---|---|
| Data type per value | FP4 E2M1 | FP4 E2M1 |
| Block size | 16 values | 32 values |
| Scale per block | FP8 E4M3 | E8M0 (power of two) |
| Additional scale | FP32 per tensor | none |
| Memory according to NVIDIA | about 3.5 times smaller than FP16, 1.8 times smaller than FP8 | - |
| Well-known example | NVFP4 checkpoints from NVIDIA and Red Hat | gpt-oss-20b and gpt-oss-120b |
The smaller blocks and the finer FP8 scale are intended to reduce rounding error compared with MXFP4. Native FP4 compute units come with the fifth-generation Tensor Cores, that is Blackwell GPUs such as the RTX PRO 4000 and RTX PRO 6000 Blackwell and the GB10 chip.
vLLM loads NVFP4 checkpoints from NVIDIA Model Optimizer and from llm-compressor (vLLM, NVIDIA Model Optimizer). At startup vLLM selects a suitable GEMM kernel. If the GPU lacks a native FP4 kernel, vLLM falls back to W4A16 via Marlin and logs a warning; the memory benefit remains, the compute benefit is lost. Which kernel is actually active is shown in the startup log and should be checked at every deployment.
MXFP4 became known mainly through gpt-oss: according to the model card, the MoE weights were quantized to MXFP4 during post-training, and all published benchmarks refer to this variant. gpt-oss-120b therefore runs on a single 80 GB GPU, gpt-oss-20b within 16 GB of memory.
On quality there are vendor figures, not an independent guarantee:
| Source | Model / size | Statement |
|---|---|---|
| NVIDIA | DeepSeek-R1-0528, FP8 vs. NVFP4 | 1 % or less deviation across seven benchmarks |
| Red Hat | 70 to 235 billion parameters | about 99 % of BF16 accuracy |
| Red Hat | about 30 billion parameters | 97 to 99 % |
| Red Hat | 7 to 14 billion parameters | about 95 to 98 % |
The trend is clear: large models and MoE models tolerate 4 bits better than small dense models.
AWQ and GPTQ: 4-bit weights with calibration
AWQ and GPTQ are not data types but methods that convert weights after training (post-training quantization), usually into 4-bit integers. Both use a small set of calibration texts to keep rounding errors low. Activations stay in 16 bits (W4A16).
| Property | GPTQ | AWQ |
|---|---|---|
| Publication | Frantar et al., ICLR 2023 | Lin et al., MLSys 2024 (Best Paper) |
| Core idea | Round weights layer by layer and compensate errors with a second-order approximation | Identify important channels from the activations and scale them before rounding |
| Bit width | 3, 4 or 8 bits | mostly 4 bits |
| Tool today | GPTQModel, llm-compressor | llm-compressor (AutoAWQ is deprecated according to vLLM) |
| vLLM kernels | Marlin, Machete | Marlin |
The AWQ authors show that protecting only about 1 % of the most important weights already reduces the quantization error considerably. According to the paper, GPTQ quantized a model with 175 billion parameters to 3 to 4 bits in about four GPU hours.
In vLLM, AWQ and GPTQ are supported on Turing, Ampere, Ada and Hopper according to the hardware table, GPTQ also on Volta. The INT4 page lists Ampere, Ada Lovelace, Hopper and Blackwell as compute platforms (vLLM, INT4 W4A16). For GPUs without FP4 compute units, AWQ and GPTQ are thus the common 4-bit formats. On Blackwell, NVFP4 is an alternative where a checkpoint is available.
GGUF: the format of llama.cpp and Ollama
GGUF is the file format of llama.cpp. It holds weights, tokenizer and metadata in one file and comes with its own quantization types. Common choices are the K-quants (such as Q4_K_M, Q5_K_M, Q6_K) and the I-quants for very low bit widths. According to llama.cpp, an importance matrix (imatrix) built from calibration texts reduces quality loss, which is measured there via perplexity and KL divergence.
Measurements from the llama.cpp documentation for Llama 3.1 8B:
| Type | Bits per weight | Size (GiB) |
|---|---|---|
| F16 | 16.0 | 14.96 |
| Q8_0 | 8.50 | 7.95 |
| Q6_K | 6.56 | 6.14 |
| Q5_K_M | 5.70 | 5.33 |
| Q4_K_M | 4.89 | 4.58 |
| IQ3_M | 3.76 | 3.52 |
| IQ2_M | 2.93 | 2.74 |
Ollama uses GGUF as its model format. According to the Ollama documentation, Ollama does not quantize GGUF models during import; quantization happens beforehand with llama-quantize. GGUF runs on CPUs, NVIDIA and AMD GPUs and Apple Silicon and can split layers between GPU and system memory. That makes it the first choice for single workstations and small installations. vLLM supports GGUF only through a separate plugin and describes the support as "highly experimental and under-optimized" (vLLM, GGUF). For vLLM, FP8, NVFP4, AWQ or GPTQ are the appropriate formats. The trade-off between the two engines is covered in vLLM vs. Ollama.
Memory footprint by example
Qwen3-32B has around 32.8 billion parameters (Apache 2.0). The memory required for the weights is the parameter count times bits per weight divided by 8:
| Format | Bits per weight (calculated) | Weights, rounded |
|---|---|---|
| BF16 | 16 | about 66 GB |
| FP8 | 8 | about 33 GB |
| GGUF Q8_0 | 8.5 | about 35 GB |
| GGUF Q4_K_M | 4.89 | about 20 GB |
| NVFP4 | 4.5 | about 18 GB |
| INT4 (AWQ/GPTQ, group 128) | slightly above 4 | about 17 GB |
These are calculated values. Real files differ because individual layers such as embeddings or the output layer often remain at higher precision.
The table shows the gradation: in BF16 the model needs a 96 GB card, in FP8 it fits there with plenty of room for the KV cache. A 24 GB card only holds it in 4 bits, and then only a few gigabytes remain for the KV cache and runtime.
Quantizing the KV cache
Besides the weights, the KV cache takes memory: for every token in the context, each layer stores keys and values. The requirement per token is:
KV per token = 2 × layers × KV heads × head dimension × bytes per value
For Qwen3-32B (64 layers, 8 KV heads, head dimension 128 according to its config.json) this gives:
| KV cache type | Per token | 32,768 tokens | 10 requests of 32,768 tokens |
|---|---|---|---|
| BF16 | 256 KiB | about 8.6 GB | about 86 GB |
| FP8 | 128 KiB | about 4.3 GB | about 43 GB |
The KV cache grows with context length and concurrency, not with the quantization of the weights. A 4-bit model with long context and many users can therefore need more memory for the cache than for the weights.
vLLM: --kv-cache-dtype fp8 (or fp8_e4m3, fp8_e5m2) stores the cache in FP8. Without calibration vLLM sets all scale factors to 1.0; calibrated scales via llm-compressor are recommended. With the Flash Attention 3 backend, queries are also computed in FP8 (vLLM, Quantized KV Cache).
Ollama: OLLAMA_KV_CACHE_TYPE allows f16 (default), q8_0 (about half the memory, according to the documentation usually without noticeable quality impact) and q4_0 (about a quarter, small to medium precision loss, more noticeable with long context). Flash Attention is required, and the setting applies globally to all models (Ollama FAQ).
Which format for which hardware
The choice of format follows from the GPU's compute capability, memory size and memory bandwidth. The hardware data comes from NVIDIA (CUDA GPUs, RTX PRO 4000 Blackwell, RTX PRO 6000 Blackwell Max-Q, DGX Spark hardware), as of October 2026:
| Hardware | Memory | Bandwidth | Compute capability | FP8 | Native FP4 |
|---|---|---|---|---|---|
| RTX PRO 4000 Blackwell | 24 GB GDDR7 ECC | 672 GB/s | 12.0 | yes | yes |
| RTX PRO 6000 Blackwell Max-Q | 96 GB GDDR7 ECC | 1,792 GB/s | 12.0 | yes | yes |
| GB10 (AI Cube) | 128 GB LPDDR5x, shared by CPU and GPU | 273 GB/s | 12.1 | yes | yes |
| For comparison: L40S, RTX 4090 (Ada) | - | - | 8.9 | yes | no |
| For comparison: Ampere (A100, RTX 30) | - | - | 8.0 / 8.6 | weights only | no |
This leads to typical starting points:
| Scenario | Obvious format | Reason |
|---|---|---|
| 24 GB, models up to about 8 billion parameters | BF16 or FP8 with FP8 KV cache | small weights, plenty of room for cache and users |
| 24 GB, models with 14 to 32 billion parameters | 4 bits: NVFP4, AWQ or GPTQ | only then do the weights fit, context limited |
| 96 GB, dense models up to about 70 billion parameters | FP8 | quality close to BF16, room for KV cache |
| 96 GB, large MoE models | NVFP4 or MXFP4 (gpt-oss) | weights only fit in 4 bits |
| AI Cube with GB10, 128 GB | 4-bit formats (MXFP4, NVFP4, GGUF Q4_K_M) | large memory but 273 GB/s bandwidth; fewer bits per weight means more tokens per second |
| Older GPUs without FP8 (Ampere) | AWQ or GPTQ, FP8 only as W8A16 | no FP8 compute units |
The boundaries are fluid and depend on the model. Which models run sensibly on 128 GB of unified memory is shown in LLM models on 128 GB unified memory; running gpt-oss-120b on the AI Cube is described in gpt-oss-120b on the AI Cube. A practical example of FP8 on two 96 GB cards is Qwen 122B with vLLM on two RTX PRO 6000. When a model is split across several GPUs is covered in Multi-GPU inference without NVLink.
Measuring quality loss instead of guessing
Vendor figures refer to public benchmarks such as MMLU, GPQA or HumanEval. They show the direction but not how a quantized model behaves with German contract clauses, structured output or tool calls. Especially with small models and at 4 bits, a comparison of your own is worthwhile.
A sound comparison has four steps:
- Set a reference. Run the model in BF16 or FP8 on the same tasks.
- Choose everyday tasks. 50 to 200 real requests with expected answers, plus cases with JSON output or tool calls if the system uses them.
- Configure identically. Same prompt, same temperature, same context length; only the format changes.
- Evaluate deviations. Do not only count hits, also check format errors, truncated answers and outliers.
For standardised benchmarks the vLLM documentation uses lm-evaluation-harness. llama.cpp ships llama-perplexity, a tool for perplexity and KL divergence against the original. For RAG systems, answer quality with sources belongs in the measurement, as described in Measuring RAG quality.
The licence also matters: a quantized variant is derived from the original model and is generally still subject to its licence terms. More on this in Open LLM licences for commercial use.
What this means for your project
Quantization is not a final polish but part of the hardware decision. Knowing which model should run in which format allows GPU size, context length and the number of concurrent users to be planned reliably. Buying the hardware first fixes the possible formats.
We set up models with vLLM or Ollama, choose the format to match the GPU, check in the startup log which kernel actually runs and compare the quantized variant against the reference using your own tasks. The basis is WZ-IT's managed GPU servers with NVIDIA RTX PRO 4000 Blackwell (24 GB GDDR7 ECC) or RTX PRO 6000 Blackwell Max-Q (96 GB GDDR7 ECC) and the AI Cube with 128 GB unified memory. Support, consulting and implementation by WZ-IT.
Model choice itself is covered in Which LLM to self-host?, the total memory requirement in Sizing GPU and VRAM for LLMs. The difference between inference and training is explained in Inference vs. training, and the inference engines are compared in vLLM, Ollama and llama.cpp compared.
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
Quantization stores the weights of a language model, and depending on the method also activations and the KV cache, with fewer bits than the original. A model in BF16 uses 16 bits per parameter, in FP8 8 bits and in 4-bit formats such as NVFP4, AWQ or GGUF Q4_K_M around 4.2 to 4.9 bits. Memory requirements fall, larger models fit on the same GPU, and because less data is read from memory, throughput often rises.
There is no best format, only one that fits the hardware, inference engine and model. FP8 is the obvious default for vLLM on GPUs from Ada Lovelace (compute capability 8.9) onwards. NVFP4 uses the FP4 Tensor Cores of Blackwell. AWQ and GPTQ are widespread 4-bit formats for vLLM on older and current GPUs. GGUF is the format of llama.cpp and Ollama. Only a test with your own tasks shows which format keeps enough quality for a given model.
Quantization changes the weights and can reduce quality; how much depends on format, model size and task. NVIDIA reports a deviation of 1 percent or less across seven benchmarks for DeepSeek-R1-0528 when moving from FP8 to NVFP4. Red Hat reports around 99 percent of BF16 accuracy for NVFP4 on models with 70 to 235 billion parameters and around 95 to 98 percent for 7 to 14 billion parameters. Small models are therefore more sensitive. These are vendor figures on public benchmarks, not a guarantee for your own tasks.
No. NVFP4 is a 4-bit floating-point format (E2M1) with an FP8 scale factor per block of 16 values and an additional FP32 factor per tensor. AWQ and GPTQ are methods that usually convert weights into 4-bit integers (INT4) using calibration data. With AWQ and GPTQ the activations typically stay in 16 bits (W4A16). On Blackwell GPUs, NVFP4 can also run the computation itself in FP4.
With limitations. Native FP4 compute units only arrived with the fifth-generation Tensor Cores in Blackwell. According to its documentation, vLLM falls back to weight-only execution (W4A16 via Marlin) on GPUs without a suitable FP4 kernel and logs a warning. The model then runs with the memory benefit but without the compute benefit of FP4.
Ollama works with GGUF, the llama.cpp format, and can also import Safetensors weights. According to the Ollama documentation, Ollama does not quantize GGUF models during import; quantization happens beforehand with llama-quantize from llama.cpp. AWQ and GPTQ checkpoints are intended for engines such as vLLM. Switching between the two worlds usually requires a different model file.
The KV cache stores the attention keys and values per token and layer. In FP8 instead of BF16 its memory requirement halves, so twice as many tokens across all concurrent requests fit into the same VRAM. In vLLM this is enabled with --kv-cache-dtype fp8. Without calibration vLLM sets all scale factors to 1.0; for higher accuracy the documentation recommends calibrated scales via llm-compressor.
Only when quantized. Qwen3-32B has around 32.8 billion parameters and needs about 66 GB in BF16 and about 33 GB in FP8. In a 4-bit format it comes to roughly 17 to 20 GB by calculation. On a 24 GB card only a few gigabytes then remain for the KV cache and runtime, which clearly limits context length and the number of concurrent requests. For more users or long context, a 96 GB card is the more suitable size.
Usually not. Many vendors publish official FP8 variants, for example Qwen3-8B-FP8, and gpt-oss was released with MoE weights that were already quantized to MXFP4 during post-training. For other formats there are tools such as llm-compressor (FP8, INT8, INT4, NVFP4, MXFP4, AWQ, GPTQ), NVIDIA Model Optimizer (FP8, NVFP4) and llama-quantize for GGUF. Quantizing yourself makes sense when a model is not available in the required variant or calibration should use your own data.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Knowledge transfer during employee transitions
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- Private ChatGPT for business
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for professional secrecy holders
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Buy or rent an AI server?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Chunking for RAG
- Hybrid search and reranking
- Contextual retrieval
- Measuring RAG quality
- Open LLM licences for commercial use
- LLM quantization
- MCP in the enterprise
- Protection against prompt injection
- Multi-GPU inference without NVLink
- Text-to-SQL
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





