WZ-IT Logo

Open LLM licences for commercial use: Apache 2.0, MIT, Llama, Qwen, Mistral

Timo WevelsiepTimo Wevelsiep•Updated: 30.09.2026

Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.

Model selected, licence clarified, operation still open? WZ-IT runs open-weight models on your own infrastructure and documents model version, quantization and licence status as part of the model approval. Models run on the AI Cube on your premises or on managed GPU servers, operated through LLM hosting. Book an introductory call · Explore LLM hosting

Open language models can be downloaded and run on your own hardware. Whether a company may also use a particular model commercially, adapt it or provide it to customers is not stated in the model name but in the licence file of the specific version. Terms can differ within a single model family, and some licences contain revenue thresholds, naming obligations or regional exclusions. This article classifies the common licence types, summarises the terms of important models and answers the questions that arise with quantization, fine-tuning and hosting for third parties. As of October 2026.

Note: This article is a technical assessment based on the published licence texts and model cards. It is not legal advice and does not replace a legal review of the individual case.

Table of contents

Open weight is not the same as open source

"Open weight" only describes availability: the trained weights of a model can be downloaded publicly. Which rights come with them is set by the accompanying licence. It decides three questions that matter to companies:

Question Where the answer is
May the model be used commercially? Licence file (LICENSE) in the model repository
Which obligations apply when distributing or providing it? Licence file, NOTICE file if present
Which uses are excluded? Licence file and, if present, an Acceptable or Prohibited Use Policy

The licence shown in the header of a Hugging Face model card is a good first indication, but not binding. If it says other, the model has its own licence whose text has to be read. What is binding is the licence file of the version in use, including the version from which a quantization or fine-tune was derived.

The licence governs usage rights to the model. Obligations under the AI Act, data protection law or professional secrecy apply independently. The regulatory framework is described in The EU AI Act for companies.

Licence types at a glance

Licence type Commercial use Key obligations Examples (as of October 2026)
Apache 2.0 yes Licence text and NOTICE on distribution, mark modified files; patent licence; no trademark rights gpt-oss, Gemma 4, Qwen3.8-27B, Mistral Small 4, Granite 4.2
MIT yes Copyright and permission notice in copies DeepSeek V4, GLM-5.3-Flash, MiMo-V2.6
Modified MIT yes, with thresholds additional naming or revenue conditions Kimi K2.6, Mistral Medium 3.5
Vendor licence with conditions yes, with conditions notice and naming obligations, usage policy, thresholds Llama 4, Qwen Community License, GLM-5.3 License
NVIDIA model licences yes include licence, attribution notice in NOTICE Nemotron 3
Non-commercial no research or evaluation use only CC BY-NC 4.0 (Jina), Qwen Research License, Mistral Non-Production License

Apache 2.0 is the most widely used permissive model licence. Section 4 requires, on distribution, a copy of the licence, change notices in modified files, retention of copyright and patent notices and, where applicable, carrying over the NOTICE file. Section 3 contains a patent licence that ends if you file patent litigation over the work. Section 6 expressly grants no trademark rights, so model names may not be used as your own product brand (Apache License 2.0).

MIT is even shorter: the copyright and permission notice must be included in all copies, with no further conditions. "Modified MIT", by contrast, means the vendor has added conditions, and these vary widely (see below).

Licences of important models

Licences of specific model versions, checked against licence file and model card, as of October 2026:

Model Licence Commercial Particulars
gpt-oss-20b, gpt-oss-120b Apache 2.0 yes Usage policy: comply with applicable law
Gemma 4 (E2B, E4B, 12B, 26B-A4B, 31B) Apache 2.0 yes Gemma 1 to 3: Gemma Terms of Use
Qwen3.8-27B, Qwen3.6-27B, Qwen3.6-35B-A3B Apache 2.0 yes -
Qwen3.8-Flash-Next Qwen Community License 1.0 restricted separate licence for model-as-a-service and AI work assistants
Qwen3.8-2.4T-A95B Qwen3.8-Max License restricted separate licence for model-as-a-service and AI work assistants above USD 50 million revenue in 12 months
Mistral Small 4, Mistral Large 3, Ministral 3, Devstral Small 2 Apache 2.0 yes -
Mistral Medium 3.5, Devstral 2 (123B) Modified MIT up to USD 20 million monthly revenue above that, commercial licence from Mistral AI required
DeepSeek V4 Pro and Flash MIT yes -
GLM-5.3-Flash, GLM-5.2 MIT yes -
GLM-5.3 GLM-5.3 License yes model-as-a-service above USD 10 billion revenue: security review by Z.AI
Kimi K2.6 Modified MIT yes above 100 million MAU or USD 20 million monthly revenue, display "Kimi K2.6" in the UI
MiMo-V2.6 MIT yes -
Granite 4.2 (3B, 8B, 30B) Apache 2.0 yes -
Apertus Apache 2.0 yes -
Nemotron 3 Nano and Super NVIDIA Nemotron Open Model License yes attribution notice on distribution
Llama 4 Scout and Maverick Llama 4 Community License with conditions "Built with Llama", 700 million MAU threshold, EU exclusion for multimodal models
Qwen2.5-72B-Instruct Qwen License yes separate licence above 100 million MAU
Qwen2.5-3B-Instruct Qwen Research License no research and evaluation only
Codestral 22B v0.1 Mistral AI Non-Production License no no production use
jina-embeddings-v3, jina-reranker-v2-base-multilingual CC BY-NC 4.0 no non-commercial only

Which of these models fit which hardware is covered in Which LLM to self-host? and GPU and VRAM sizing for LLMs.

Llama: community licence with an EU exclusion

The Llama 4 Community License (effective 5 April 2025) permits use, modification and distribution, but attaches conditions:

Provision Content
Distribution or making available (1.b.i) Include a copy of the licence; prominently display "Built with Llama" on a website, UI or documentation
Derived models (1.b.i) Anyone using Llama or Llama outputs to train or improve a model that is made available puts "Llama" at the beginning of the model name
Attribution notice (1.b.iii) NOTICE file with "Llama 4 is licensed under the Llama 4 Community License, Copyright © Meta Platforms, Inc. All Rights Reserved."
Usage policy (1.b.iv) Acceptable Use Policy is incorporated into the agreement
User threshold (2) more than 700 million monthly active users on the release date: separate licence from Meta
Governing law (7) Laws of the State of California, exclusive jurisdiction of California courts

For companies in the EU, one rule matters more than the user threshold. The Llama 4 Acceptable Use Policy does not grant the rights under section 1(a) for multimodal models included in Llama 4 to individuals domiciled in, or companies with their principal place of business in, the EU. End users of a product or service incorporating such a model are exempt. According to the model card, the Llama 4 models are natively multimodal. The same rule already appears in the Llama 3.2 usage policy for its multimodal models.

For a company based in Germany or elsewhere in the EU that wants to run Llama 4 itself, this is an exclusion that has to be assessed legally before deployment. Comparable models under Apache 2.0 or MIT avoid the question.

Qwen, Mistral and GLM: one family, several licences

With several vendors, the licence depends on the model size or the release. A blanket approval such as "Qwen is Apache 2.0" or "Mistral is Apache 2.0" is therefore not reliable.

Qwen. The mid-sized models Qwen3.8-27B and Qwen3.6 are under Apache 2.0. Qwen3.8-Flash-Next is under the Qwen Community License 1.0: if the licensee or an affiliate runs a model-as-a-service or AI work assistant business, it needs a separate licence before any commercial use. Internal use is exempt as long as neither the model nor its outputs or capabilities are made available to third parties. Model-as-a-service here means third-party access to inference or fine-tuning with meaningful control over inputs, parameters or training data. The Qwen3.8-Max License for the 2.4-trillion-parameter model applies the same requirement only above USD 50 million revenue in twelve months. Both licences additionally require the model name to be displayed prominently if the product has more than 100 million monthly active users or USD 20 million monthly revenue.

Mistral. Mistral Small 4, Mistral Large 3, Ministral 3 and Devstral Small 2 are under Apache 2.0. Mistral Medium 3.5 and Devstral 2 (123B) are under a Modified MIT License with a hard limit: if the company's global consolidated monthly revenue exceeded USD 20 million in the preceding month, no rights under the licence may be exercised. This applies to the model and all derivatives, including internal use. Older models such as Codestral 22B v0.1 are under the Mistral AI Non-Production License and are not cleared for production use.

GLM. GLM-5.3-Flash and GLM-5.2 are under MIT. The larger GLM-5.3 has its own licence: if the licensee runs a model-as-a-service business and revenue exceeds USD 10 billion over twelve months, a security review by Z.AI is required before commercial use. For mid-sized companies this has practically no effect, but it changes the licence category.

NVIDIA licences and non-commercial models

NVIDIA uses several model licences that differ in one point:

Licence Commercial Distribution Guardrail clause
NVIDIA Nemotron Open Model License yes include licence, notice "Licensed by NVIDIA Corporation under the NVIDIA Nemotron Model License" no
NVIDIA Open Model Agreement yes notice "Licensed by NVIDIA Corporation under the NVIDIA Open Model Agreement" no
NVIDIA Open Model License yes notice "Licensed by NVIDIA Corporation under the NVIDIA Open Model License" yes
OpenMDW 1.1 (e.g. Nemotron 3.5 Lightning) yes retain licence and notices of origin no

The older NVIDIA Open Model License terminates the rights automatically if technical limitations or safety guardrails of the model are bypassed or weakened without a substantially similar replacement. This matters for fine-tuning that changes safety behaviour. All three NVIDIA licences claim no ownership of model outputs and terminate if you file patent or copyright litigation over the work.

Non-commercial licences rule out company use without a separate agreement. This mainly affects auxiliary models in RAG systems: jina-embeddings-v3 and jina-reranker-v2-base-multilingual are under CC BY-NC 4.0. Embedding and reranker models are therefore checked just like the language model. Permissively licensed alternatives are listed in Hybrid search and reranking.

Practical questions: quantization, fine-tuning, hosting, outputs

Scenario Apache 2.0 / MIT Llama 4 Community License Other vendor licences
Internal use, unchanged permitted permitted, observe AUP; EU exclusion for multimodal observe revenue limit of Mistral Modified MIT
Quantized version, internal permitted permitted as original model
Distributing a quantized version include licence and NOTICE, mark changes include licence, "Built with Llama", NOTICE licence and notice obligations of the respective licence
Fine-tune, internal permitted permitted as original model
Publishing a fine-tune licence and NOTICE, mark changes name starts with "Llama" e.g. NVIDIA: attribution notice
Model as a service for customers permitted display "Built with Llama" Qwen Community License: separate licence; GLM-5.3: review above USD 10 billion
Outputs used to train other models not addressed separately model made available carries "Llama" in its name Gemma Terms of Use (up to Gemma 3): distilled model is a Model Derivative

Quantization. A model in FP8, AWQ, GPTQ or GGUF is a modified version of the weights. The licence stays the same; quantization does not remove any conditions. For community quantizations on Hugging Face, the base model is therefore checked, not only the information on the quantizer's page. The technical formats are explained in LLM quantization.

Fine-tuning. A fine-tune, for example a LoRA adapter, is a derivative work of the base model. As long as it stays within the company, Apache 2.0 and MIT add no further obligations. If it is published or passed on to customers, the distribution obligations apply, and for Llama also the naming rule.

Hosting for third parties. Apache 2.0 and MIT tie their obligations to the distribution of copies. Providing a model only via an API does not distribute any weights. Vendor licences, by contrast, sometimes address this case explicitly: Llama requires "Built with Llama" for services made available as well, and the Qwen Community License and the GLM-5.3 License address model-as-a-service directly.

Outputs. None of the licences reviewed claims rights to the generated text. Google states in the Gemma Terms of Use that it claims no rights in outputs, as does NVIDIA in its model licences. Restrictions concern the use of outputs to train other models. Whether copyright arises in the outputs themselves is not a question of the model licence.

Licence review as part of model approval

A licence review is not a one-off step but part of the model approval. A fixed scheme works well:

  1. Pin the specific version: repository, model ID and revision (commit hash) of the base model, not just the family name.
  2. Read the licence file: LICENSE and, where present, NOTICE, USAGE_POLICY or an Acceptable Use Policy. The entry in the model card header is not sufficient.
  3. Check the derivation chain: for quantizations, fine-tunes and distillations, identify the original model and apply its licence.
  4. Assign the deployment scenario: internal, customer product, service for third parties, distribution of weights. Check the obligations from the table above for each scenario.
  5. Check thresholds: revenue and user thresholds (Mistral Modified MIT, Qwen, Kimi, Llama) against your own figures, including affiliates.
  6. Include auxiliary models: treat embedding, reranker, vision and speech recognition models the same way.
  7. Document the result: record licence status with date, reviewer and deployment scenario in the model register; review again when the model changes.

The same model register serves as the basis for updates and rollbacks in operation, because it records which version was approved with which licence status.

What this means for your project

For most company use cases, capable models are available under Apache 2.0 or MIT, for example gpt-oss, Gemma 4, Qwen3.8-27B, Mistral Small 4, DeepSeek V4 or Granite 4.2. With these licences, internal use, quantization and fine-tuning remain free of additional obligations. Vendor licences with revenue thresholds, naming rules or regional exclusions, on the other hand, require a legal assessment before deployment.

WZ-IT runs open-weight models with vLLM or Ollama on the AI Cube or on managed GPU servers with NVIDIA RTX PRO 4000 Blackwell (24 GB) or RTX PRO 6000 Blackwell Max-Q (96 GB). In managed AI operations, model version, quantization and licence status are documented per model and carried forward with updates. The legal assessment remains the responsibility of the company and its legal counsel. Support, consulting and implementation by WZ-IT.

Which models suit which task is described in Which LLM to self-host?, and a comparison of current model families is provided in Llama 4 vs. Qwen 3.5 vs. DeepSeek V4. Anyone who prefers to use models via European providers rather than running them in-house will find an overview in European LLM APIs compared.

Sources

Rather have it operated?

You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.

Enquiry

Assess local AI for your use case

Start with the AI Cube or have us assess a custom AI platform, knowledge connection, or integration.

How should we get back to you?

Frequently Asked Questions

Answers to the most important questions

In many cases yes, but not across the board. Models under Apache 2.0 or MIT, such as gpt-oss, Gemma 4, Qwen3.8-27B, DeepSeek V4 or Granite 4.2, allow commercial use with few obligations. Other models come with their own licences that contain revenue or user thresholds, naming obligations or exclusions. What counts is always the licence file of the specific model version.

No. Open weight only means the weights can be downloaded. Usage rights are set by the accompanying licence, which ranges from Apache 2.0 to licences that rule out commercial use, such as the Qwen Research License or CC BY-NC 4.0. The Llama Community License is not an open source licence in the classic sense either, because it contains usage conditions and exclusions.

In practice, no. Under section 2 of the Llama 4 Community License, only licensees whose products had more than 700 million monthly active users in the preceding calendar month on the Llama 4 release date need a separate licence from Meta. More relevant for mid-sized companies are the 'Built with Llama' notice, the naming rule for derived models and the Acceptable Use Policy.

Only with restrictions. The Llama 4 Acceptable Use Policy does not grant the rights for multimodal models included in Llama 4 to individuals domiciled in, or companies with their principal place of business in, the EU. According to the model card, Llama 4 Scout and Maverick are natively multimodal. End users of a product that incorporates such a model are exempt. For self-hosting by an EU company this is an exclusion that legal counsel should assess.

No. Qwen3.8-27B and the Qwen3.6 models are under Apache 2.0. Qwen3.8-Flash-Next is under the Qwen Community License 1.0, which requires a separate licence for model-as-a-service and AI work assistant businesses. Older models such as Qwen2.5-72B-Instruct (Qwen License) or Qwen2.5-3B-Instruct (non-commercial only) have different terms again. The licence is checked per model, not per family.

Gemma 4 is. Google released Gemma 4 under Apache 2.0 on 2 April 2026. Gemma 1 to 3, Gemma 3n and variants such as EmbeddingGemma or PaliGemma remain under the Gemma Terms of Use with a Prohibited Use Policy and distribution obligations. Anyone running Gemma 3 is therefore checking different terms than for Gemma 4.

For the Apache 2.0 models such as Mistral Small 4, Mistral Large 3, Ministral 3 or Devstral Small 2, yes. Mistral Medium 3.5 and Devstral 2 (123B) are under a Modified MIT License: companies whose global consolidated monthly revenue exceeded 20 million US dollars in the preceding month may not exercise any rights under it and need a commercial licence from Mistral AI. This includes internal use.

Quantization changes the model and therefore creates a derivative work. For internal use, Apache 2.0 and MIT add no further obligations. Anyone distributing quantized weights must meet the obligations of the original licence, for Apache 2.0 for example including the licence text, carrying over the NOTICE file and marking modified files. Quantization does not change the licence.

Under Apache 2.0 and MIT, yes. Some licences address exactly this case: the Qwen Community License 1.0 requires a separate licence for model-as-a-service, the GLM-5.3 License requires a security review by Z.AI above 10 billion US dollars in annual revenue, and the Llama licence requires the 'Built with Llama' notice for services made available.

That depends on the licence. Apache 2.0 and MIT do not address outputs separately. The Llama 4 Community License requires that a model trained with Llama outputs and made available carries 'Llama' at the beginning of its name. The Gemma Terms of Use (up to Gemma 3) count models created by distillation as Model Derivatives, but not the outputs themselves.

Not without a separate agreement. According to their model cards, jina-embeddings-v3 and jina-reranker-v2-base-multilingual are licensed under CC BY-NC 4.0, i.e. for non-commercial purposes only. For RAG systems in companies, models under Apache 2.0 or MIT, such as BGE-M3 or the Qwen3 embedding and reranker models, are suitable instead.

No. The article is a technical assessment based on the published licence texts and model cards, as of October 2026, and not legal advice. Licences can change with new model versions. A binding assessment, for example when passing models on to customers or building products on them, requires a legal review.

More on Local AI for Business

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.