WZ-IT Logo
Llama hosting

Llama Hosting in Germany: Run Llama 3.3, Understand Llama 4 in the EU

You want to use a Llama model on your own infrastructure. We explain which Llama version a company in the EU can use under the licence, and run Llama 3.3 70B on a dedicated GPU server in Germany.

Operated in a German data centreOperation under a DPAvLLM and OpenAI-compatible API

Companies worldwide trust WZ-IT

Reviews

Companies worldwide trust WZ-IT

Stadtwerke BrühlDGHO e.V.ABCO Water SystemsGolem.deEVADXBnextGYMAInergyml&sOdiseo SolutionsAnnotaARGESweetConnect GmbH
Starting point

Why run Llama yourself

Llama is Meta's open-weight family. The latest generation, Llama 4 (Scout and Maverick, April 2025), consists of multimodal mixture-of-experts models. The Llama 4 Acceptable Use Policy does not grant the rights under Section 1(a) of the licence for these multimodal models to companies with their principal place of business in the EU. The only exception is end users of a product or service that incorporates the models. We therefore do not offer Llama 4.

Llama 3.3 70B Instruct can be operated: a text-only model from December 2024 under the Llama 3.3 Community License. Since April 2025, Meta has not published new model weights in the meta-llama Hugging Face account. If you are looking for a current model, Qwen, gpt-oss, Gemma and Mistral offer families with newer checkpoints under Apache 2.0.

  • Data stays in Germany

    The model runs on a dedicated server in a German data centre. Under a data processing agreement, with no data path to a model API.

  • Predictable costs

    A fixed monthly price per server tier instead of billing per consumed token. More requests do not increase the invoice.

  • Model version under control

    Checkpoint and quantisation stay until you agree to a change. No silent model update by a provider.

Model facts

Llama checkpoints at a glance

Vendor specifications and weight file size per format, with source.

CheckpointReleasedParametersContext (vendor)Weights per format
meta-llama/Llama-3.3-70B-Instruct12/202470.6B (dense, text only)131,072 tokens
meta-llama/Llama-4-Scout-17B-16E-Instruct04/2025109B total / 17B active (MoE, multimodal)10M tokens
meta-llama/Llama-4-Maverick-17B-128E-Instruct04/2025400B total / 17B active (MoE, multimodal)1M tokens

GB = 10⁹ bytes, sum of the weight files in the listed Hugging Face repository. Operation additionally needs memory for KV cache, runtime and image processing where applicable. As of October 2026.

Licence

Llama: licence and commercial use

Licence
Llama 3.3 Community License + Acceptable Use Policy (Licence text, Llama 3.3 Acceptable Use Policy)
Commercial use
Yes, for Llama 3.3 70B Instruct. Anyone whose products or services had more than 700 million monthly active users on the release date needs a separate licence from Meta.
Obligations and thresholds
Anyone making the model or a service with it available includes the licence and prominently displays “Built with Llama” on a website, user interface or in the documentation. The Acceptable Use Policy must be followed. Its EU clause concerns multimodal models; Llama 3.3 70B Instruct is a text-only model.
Checked
As of October 2026

Not offered

Llama 4 Scout and Llama 4 Maverick are multimodal models. The Llama 4 Acceptable Use Policy does not grant the rights under Section 1(a) of the Llama 4 Community License to individuals domiciled in, or companies with their principal place of business in, the EU. The exception only covers end users of a product or service that incorporates the models. As a company based in Germany, we therefore do not offer Llama 4. Licence text

Technical classification, not legal advice. The licence text of the deployed model version is authoritative.

Open LLM licences compared
Server tier

Minimum recommended server tier for Llama

The recommendation follows from the weight size plus headroom for context and runtime.

ModelFormatWeightsMinimum recommendedNote
Llama 3.3 70BFP872.7 GBManaged GPU Server 96On 96 GB, around 23 GB remain for context and runtime after the weights. This limits context length and concurrent requests.
Llama 3.3 70BBF16141.1 GBManaged GPU Server 192
Llama 4 ScoutBF16217.3 GBNone of our tiersNot offered: no rights under Section 1(a) for companies with their principal place of business in the EU.
Llama 4 MaverickFP8416.8 GBNone of our tiersNot offered: EU exclusion in the licence; the weights also exceed our largest tier of 384 GB.

Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

Meta publishes Llama 3.3 70B in BF16 only. The FP8 size comes from Red Hat AI's checkpoint in compressed-tensors format, licensed under the Llama 3.3 Community License. NVIDIA's FP8 and NVFP4 checkpoints are additionally subject to the NVIDIA Open Model License.

The Managed GPU Server tiers

Managed GPU Server 24

1 × RTX PRO 4000 Blackwell, 24 GB GPU memory

€699 excl. VAT / month, cancellable monthly

€499 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 96

1 × RTX PRO 6000 Blackwell Max-Q, 96 GB GPU memory

€1,799 excl. VAT / month, cancellable monthly

€999 excl. VAT one-time setup

View configuration

Relevant for this family

Managed GPU Server 192

2 × RTX PRO 6000 Blackwell Max-Q, 192 GB GPU memory

Price and term on request

View configuration

Managed GPU Server 384

4 × RTX PRO 6000 Blackwell Max-Q, 384 GB GPU memory

Price and term on request

View configuration

The Managed GPU Server 288 with three GPUs is intended for several models side by side, because vLLM only splits a model when the attention heads are divisible by the number of GPUs.

vLLM

Running Llama with vLLM

Architecture, quantisation and distribution across several GPUs.

Architecture

Llama 3.3 70B uses LlamaForCausalLM, one of the longest-supported architectures in vLLM. No additional package is required for serving.

Quantisation

Meta publishes BF16 weights only. For the Managed GPU Server 96 we use an FP8 checkpoint in compressed-tensors format. Before the proposal we check the origin, licence and answer quality of the quantised checkpoint.

Tensor parallelism

With 64 attention heads and 8 KV heads, Llama 3.3 70B can be split across two or four GPUs. For BF16 at around 141 GB we recommend at least two RTX PRO 6000.

Context

Llama 3.3 supports 131,072 tokens. In FP8 on a single 96 GB GPU little memory remains for the KV cache, so the configured context length is shorter there. We set the value based on your requests and state it in the proposal.

Operated by WZ-IT

What operation includes

The Managed GPU Server is an operated model environment, not an empty server.

Model and vLLM

An agreed model in the agreed quantisation, served through a managed vLLM inference layer.

OpenAI-compatible API

Applications and coding clients connect to the server via base URL and API key.

Open WebUI

A managed chat interface for teams using the model without their own application.

24/7 monitoring

Host, GPU, vLLM and Open WebUI are monitored proactively; incidents are handled according to the service level.

Updates with CVE assessment

Operating system, drivers, vLLM and Open WebUI are reviewed and updated in a controlled way. Model changes only by agreement.

DPA and documentation

Data processing agreement, documented configuration and a personal point of contact.

Scope, prices and multi-GPU servers are on the product page.

View Managed GPU Server

Model assessment

Have your model and server tier assessed

Name the model, context length and concurrent requests. We check checkpoint, licence and server tier before the proposal.

Which server tier should be assessed?

The selection is non-binding. We confirm model fit, licence and term before a proposal.

How should we get back to you?

We usually respond within one business day. Please do not send access credentials yet.

Frequently asked questions about Llama hosting

Licence, server tier, data path and operation

Not on the basis of the Llama 4 Community License. The Acceptable Use Policy does not grant the rights under Section 1(a) for the multimodal Llama 4 models to companies with their principal place of business in the EU. End users of a product or service that incorporates the models are exempt. In our assessment, this exception does not cover running the models on your own infrastructure. We therefore do not offer Llama 4.

Llama 3.3 70B Instruct. In FP8 the minimum recommended tier is the Managed GPU Server 96, with limited room for context. In BF16 the minimum recommended tier is the Managed GPU Server 192. Which tier suits your use case depends on context length and concurrent requests. We check this before the proposal.

The Llama 3.3 Acceptable Use Policy contains the same clause, referring to multimodal models in Llama 3.3. Llama 3.3 70B Instruct only accepts and produces text. The licence text of the deployed version is authoritative; the classification on this page is not legal advice.

When redistributing or making it available: include the licence and prominently display “Built with Llama”, follow the Acceptable Use Policy and start the name of derived models with “Llama”. Companies with more than 700 million monthly active users on the release date need their own licence from Meta.

Llama 3.3 70B dates from December 2024, and according to Meta its training data has a knowledge cutoff of December 2023. No new Llama weights have been released since April 2025. If existing applications are tuned to Llama 3.3, you can keep running it. For new projects we also compare Qwen3.8-27B, gpt-oss-120b, Gemma 4 and Mistral Small 4.

Meta lists German among the eight officially supported languages of Llama 3.3. We test answers with your own examples before the model goes into production.

No. The weights run on your dedicated server in a German data centre, with no connection to Meta. Requests, documents and answers stay on the server and in the systems you connect.

The entry point for Llama is the Managed GPU Server 96, the minimum recommended tier for Llama 3.3 70B: €1,799 excl. VAT / month plus €999 excl. VAT one-time setup, including the dedicated server, vLLM, Open WebUI and managed operation by WZ-IT. Which tier fits your use depends on context length and concurrent requests. We check this before the proposal.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.