You want to use a Llama model on your own infrastructure. We explain which Llama version a company in the EU can use under the licence, and run Llama 3.3 70B on a dedicated GPU server in Germany.
Companies worldwide trust WZ-IT
Llama is Meta's open-weight family. The latest generation, Llama 4 (Scout and Maverick, April 2025), consists of multimodal mixture-of-experts models. The Llama 4 Acceptable Use Policy does not grant the rights under Section 1(a) of the licence for these multimodal models to companies with their principal place of business in the EU. The only exception is end users of a product or service that incorporates the models. We therefore do not offer Llama 4.
Llama 3.3 70B Instruct can be operated: a text-only model from December 2024 under the Llama 3.3 Community License. Since April 2025, Meta has not published new model weights in the meta-llama Hugging Face account. If you are looking for a current model, Qwen, gpt-oss, Gemma and Mistral offer families with newer checkpoints under Apache 2.0.
The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.
A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.
Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.
Vendor specifications and weight file size per format, with source.
| Checkpoint | Released | Parameters | Context (vendor) | Weights per format |
|---|---|---|---|---|
| meta-llama/Llama-3.3-70B-Instruct | 12/2024 | 70.6B (dense, text only) | 131,072 tokens | |
| meta-llama/Llama-4-Scout-17B-16E-Instruct | 04/2025 | 109B total / 17B active (MoE, multimodal) | 10M tokens |
|
| meta-llama/Llama-4-Maverick-17B-128E-Instruct | 04/2025 | 400B total / 17B active (MoE, multimodal) | 1M tokens |
GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.
Llama 4 Scout and Llama 4 Maverick are multimodal models. The Llama 4 Acceptable Use Policy does not grant the rights under Section 1(a) of the Llama 4 Community License to individuals domiciled in, or companies with their principal place of business in, the EU. The exception only covers end users of a product or service that incorporates the models. As a company based in Germany, we therefore do not offer Llama 4. Licence text
Technical classification, not legal advice. The licence text of the deployed model version is authoritative.
Open LLM licences comparedThe recommendation follows from the weight size plus headroom for context and runtime.
| Model | Format | Weights | Minimum recommended | Note |
|---|---|---|---|---|
| Llama 3.3 70B | FP8 | 72.7 GB | Managed GPU Server 96 | On 96 GB, around 23 GB remain for context and runtime after the weights. This limits context length and concurrent requests. |
| Llama 3.3 70B | BF16 | 141.1 GB | Managed GPU Server 192 | |
| Llama 4 Scout | BF16 | 217.3 GB | None of our tiers | Not offered: no rights under Section 1(a) for companies with their principal place of business in the EU. |
| Llama 4 Maverick | FP8 | 416.8 GB | None of our tiers | Not offered: EU exclusion in the licence; the weights also exceed our largest tier of 384 GB. |
Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.
Meta publishes Llama 3.3 70B in BF16 only. The FP8 size comes from Red Hat AI's checkpoint in compressed-tensors format, licensed under the Llama 3.3 Community License. NVIDIA's FP8 and NVFP4 checkpoints are additionally subject to the NVIDIA Open Model License.
1 × RTX PRO 4000 Blackwell, 24 GB GPU memory
€699 excl. VAT / month, cancellable monthly
€499 excl. VAT one-time setup
View configurationRelevant for this family
1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory
€1,799 excl. VAT / month, cancellable monthly
€999 excl. VAT one-time setup
View configurationRelevant for this family
2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory
Price and term on request
View configuration4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory
Price and term on request
View configurationThe Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.
Architecture, quantisation and distribution across several GPUs.
Llama 3.3 70B uses LlamaForCausalLM, one of the longest-supported architectures in vLLM. No additional package is required for serving.
Meta publishes BF16 weights only. For the Managed GPU Server 96 we use an FP8 checkpoint in compressed-tensors format. Before the proposal we check the origin, licence and answer quality of the quantised checkpoint.
With 64 attention heads and 8 KV heads, Llama 3.3 70B can be split across two or four GPUs. For BF16 at around 141 GB we recommend at least two RTX PRO 6000.
Llama 3.3 supports 131,072 tokens. In FP8 on a single 96 GB GPU little memory remains for the KV cache, so the configured context length is shorter there. We set the value based on your requests and state it in the proposal.
The Managed GPU Server is an operated model environment, not an empty server.
An agreed model in the agreed quantisation, served through a managed vLLM inference layer.
Applications and coding clients connect to the server via base URL and API key.
A managed chat interface for teams using the model without their own application.
Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.
Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.
Data processing agreement, documented configuration and a personal point of contact.
Scope, prices and multi-GPU servers are on the product page.
View Managed GPU ServerQwen hosting
Qwen hosting in Germany: Qwen3.8-27B, Qwen3.5-122B and Qwen3-Coder-Next on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familygpt-oss hosting
gpt-oss hosting in Germany: OpenAI's gpt-oss-120b and gpt-oss-20b on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyGemma hosting
Gemma hosting in Germany: Google's Gemma 4 31B, 26B A4B and 12B on a dedicated GPU with vLLM, an OpenAI-compatible API and operation under a DPA.
View familyMistral hosting
Mistral hosting in Germany: Mistral Small 4, Devstral Small 2 and Mistral Medium 3.5 on a dedicated GPU, with vLLM, licence review and operation under a DPA.
View familyModel assessment
Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.
Licence, server tier, data path and operation
Not on the basis of the Llama 4 Community License. The Acceptable Use Policy does not grant the rights under Section 1(a) for the multimodal Llama 4 models to companies with their principal place of business in the EU. End users of a product or service that incorporates the models are exempt. In our assessment, this exception does not cover running the models on your own infrastructure. We therefore do not offer Llama 4.
Llama 3.3 70B Instruct. In FP8 the minimum recommended tier is the Managed GPU Server 96, with limited room for context. In BF16 the minimum recommended tier is the Managed GPU Server 192. Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.
The Llama 3.3 Acceptable Use Policy contains the same clause, referring to multimodal models in Llama 3.3. Llama 3.3 70B Instruct only accepts and produces text. The licence text of the deployed version is authoritative; the classification on this page is not legal advice.
When redistributing or making it available: include the licence and prominently display “Built with Llama”, follow the Acceptable Use Policy and start the name of derived models with “Llama”. Companies with more than 700 million monthly active users on the release date need their own licence from Meta.
Llama 3.3 70B dates from December 2024, and according to Meta its training data has a knowledge cutoff of December 2023. No new Llama weights have been released since April 2025. If existing applications are tuned to Llama 3.3, you can keep running it. For new projects we also compare Qwen3.8-27B, gpt-oss-120b, Gemma 4 and Mistral Small 4.
Meta lists German among the eight officially supported languages of Llama 3.3. We test answers with your own examples before the model goes into production.
No. The weights run on your dedicated server in a German data centre, with no connection to Meta. Requests, documents and answers stay on the server and in the systems you connect.
The entry point for Llama is the Managed GPU Server 96, the minimum recommended tier for Llama 3.3 70B: €1,799 excl. VAT / month plus €999 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.