WZ-IT Logo

vLLM, Ollama or llama.cpp: which inference server for business use

Timo Wevelsiep
Timo Wevelsiep
••Updated: 01.10.2026
#vLLM #Ollama #llamacpp #SGLang #LocalAI #Inference

Editorial note: The information in this article was compiled to the best of our knowledge at the time of publication. Technical details, prices, versions, licensing terms, and external content may change. Please verify the information provided independently, particularly before making business-critical or security-related decisions. This article does not replace individual professional, legal, or tax advice.

vLLM, Ollama or llama.cpp: which inference server for business use

Local AI for several people at once?

The AI Cube ships with vLLM and Open WebUI preinstalled and generates up to ten answers at the same time, for teams of five to twenty people: from €6,490 excl. VAT one-time plus AI Cube Care from €349.90 excl. VAT per month. For more load we size GPU servers from €699 excl. VAT per month against your actual usage or provide managed LLM hosting.

Explore AI Cube · Clarify your sizing with us

The question of the right inference server is almost always framed as a speed question and almost always answered wrongly, because one figure is missing: how many people work at the same time.

For a single user, the three common tools sit close together. At ten concurrent requests they separate by factors, not percentages. Anyone choosing on single-user measurements and then putting a team on it has bought something that does not hold up in daily use.

This article assesses the three along that figure. All figures from the repositories, as of 24 August 2026. Since October 2026 a section on SGLang, the fourth widely used serving system, has been added (as of 1 October 2026).

Table of contents

Three tools, three levels

The most common mistake in comparisons is placing the three side by side as if they were interchangeable. They sit at different levels:

llama.cpp is the engine. It runs the model, works on CPU and GPU, and is the foundation much else builds on.

Ollama is the usability layer. It uses llama.cpp and adds convenient management on top: pull models, start them, address them. Comparing Ollama with llama.cpp therefore compares a wrapper with its contents.

vLLM is a standalone serving system. It does not build on llama.cpp but takes a different approach designed for GPU operation with many concurrent requests.

It follows that the real decision lies between the llama.cpp branch and vLLM. Whether you address llama.cpp directly or through Ollama is a question of convenience.

The three at a glance

Figures from the public repositories, as of 24 August 2026.

Tool GitHub stars Licence Level Designed for
Ollama 179,309 MIT usability layer single seat, quick start
llama.cpp 125,380 MIT engine CPU and GPU, maximum control
vLLM 89,860 Apache-2.0 serving system GPU, many concurrent requests

All three are unrestricted open source and actively developed. The star counts reflect visibility, not suitability: Ollama is the best known because it is the easiest entry point, not because it is the best choice for production.

The decisive difference: concurrent users

This is where the tools separate, and not gradually.

Published comparisons cite factors of ten and more for vLLM under concurrency against the llama.cpp branch. The reason is not a better implementation of the same thing but a different technique: vLLM batches requests arriving at the same time and processes them together rather than sequentially.

With one user that advantage largely disappears. There, llama.cpp used directly is even slightly faster than Ollama, because Ollama as a wrapper adds overhead, on the order of a low double-digit percentage.

Practical consequence: the jump from one to ten users is not a scaling problem solved with more hardware. It is an architectural question. A setup that works well for one person does not become good for a team through a larger graphics card but through a different serving approach.

The question before any hardware purchase is therefore: how many people actually work at the same time? Not how many should get access. Those two numbers usually differ by an order of magnitude.

Several people at once, but no GPU of your own?

The Managed GPU Server runs vLLM in a German data centre, with an agreed model, Open WebUI and operation by WZ-IT, from €699 excl. VAT per month and cancellable monthly.

Explore Managed GPU Server · Explore LLM hosting

llama.cpp

The engine beneath much else, under MIT, with 125,380 stars. Runs on CPU, on GPU and on Apple silicon, and is the right choice wherever no GPU is available.

Strength: runs practically everywhere, the smallest resource footprint, the greatest control over settings. Anyone working with embedded systems or CPU-only machines ends up here.

Limit: no management layer, no model management, no multi-user operation in a meaningful sense.

Ollama

The most widely adopted of the three at 179,309 stars, under MIT. Installation and a first model take minutes, and the handling borrows from container tooling.

Strength: by far the easiest entry point. For development, demos and the question "is a local model any good for our case at all", Ollama is the right tool.

Limit: it remains a layer over llama.cpp and inherits its behaviour under concurrent requests. For a team working in parallel it is the wrong foundation, and that only becomes apparent in operation.

vLLM

Under Apache-2.0, 89,860 stars, built for GPU serving. Offers an OpenAI-compatible API, so applications usually connect without modification.

Strength: the only one of the three designed for many concurrent requests. That is where the factors sit, not the percentages.

Limit: practically requires a GPU, is more involved to set up than Ollama, and is oversized for a single user. Anyone just experimenting does not start here.

SGLang as an alternative to vLLM

SGLang sits on the same level as vLLM: a standalone serving system for GPU operation with many concurrent requests. It is hosted by LMSYS, a non-profit open-source organisation, and offers an OpenAI-compatible interface (SGLang documentation).

Attribute SGLang, as of 1 October 2026
Licence Apache-2.0
GitHub stars 36,685
Current version v0.5.20, 18 September 2026
Release cadence roughly every two weeks (v0.5.17 to v0.5.20 between 8 August and 18 September 2026)
Core technique RadixAttention: reuse of shared prefixes in the KV cache
Interface OpenAI-compatible
Hardware per README NVIDIA (including H100, B200, RTX PRO 6000, DGX Spark), AMD Instinct, Google TPU, Intel, Apple Silicon, Huawei Ascend

What RadixAttention does. SGLang stores the KV cache of already processed tokens in a prefix tree. When several requests start with the same text, such as the same system prompt, the same tool descriptions of an agent or the same conversation history so far, that part is not recomputed. The SGLang paper from 2023 reports up to 6.4 times the throughput of the systems of that time, measured on such workloads. The figure is not a current comparison with vLLM: vLLM now also has automatic prefix caching.

Where SGLang is strong. Agent workloads with long, recurring prefixes, multi-turn conversations and large deployments across several GPUs. New models often arrive in both systems promptly; the v0.5.20 release notes list, for example, GLM-5.3-Flash and Qwen3.8-Flash-Next, plus optimisations for DeepSeek-V4 on RTX PRO 6000.

What to watch in operation.

  • Version changes. The two-week cadence also brings breaking changes. v0.5.20, among other things, removed prefill context parallelism v1, and the Responses API now only stores results with --enable-response-store (release notes).
  • Driver level. v0.5.19 was the last version with wheels and images for CUDA 12. Updating to v0.5.20 requires a matching driver and CUDA level.
  • Blanket speed claims do not hold. Whether SGLang or vLLM is faster depends on model, quantization, hardware, version and above all on the share of shared prefixes. Only a test with the target model and your own load is reliable.

For the question of this article SGLang changes little: the line still runs between the llama.cpp branch for single seats and a serving system for many concurrent users. SGLang is a second option on the serving side, not a substitute for asking about actual concurrency.

What we ship in the AI Cube and why

In the AI Cube we ship vLLM together with Open WebUI preinstalled. That decision follows directly from the device's purpose: it sits inside a company and is used by several people.

With Ollama, setup would have been simpler and the device worse in daily use. The experience behind that is unremarkable: a setup that convinces in a test by one person disappoints at five concurrent requests, and then the purchase is questioned rather than the software choice.

Open WebUI sits on top as the interface, because users expect a chat interface with permissions, knowledge areas and history, not a programming interface. How that connects to knowledge sources belongs to setup rather than to the inference server.

Which tool for which situation

Starting point Obvious choice Reason
First trial, is local AI viable at all Ollama minutes to a first result
One user, ongoing use, GPU available Ollama or llama.cpp vLLM's advantage does not apply
No GPU, CPU only or embedded llama.cpp the only one that runs sensibly there
Team from around five concurrent users vLLM throughput higher by factors under concurrency
Agents with long, recurring prefixes vLLM or SGLang both with prefix caching; choose after a test with your own load
Production application with user management vLLM plus Open WebUI serving and interface separated

What a throughput comparison does not answer

Does the model fit in memory at all? The choice of serving system is secondary to whether model and context length fit into available GPU memory. An oversized model is not made fast by any serving approach.

How long are the contexts? Long inputs, for instance in document search, consume memory per concurrent request. The number of users that can be served in parallel drops considerably, and that is regularly overlooked in sizing.

Who updates models and the serving layer? Both move quickly. A setup left untouched for a year keeps running but loses access to newer models.

What happens if the GPU fails? A local AI system without a fallback is entirely down on a hardware fault. Whether that is acceptable belongs settled before the purchase.

Still unsure which tool fits? The AI opportunity check tests which approach holds up on a concrete process and ends with a pilot recommendation. From €3,900 excl. VAT.

How we approach this at WZ-IT

We ask about the number of concurrent users before choosing a tool, and we insist, because the first answer is almost always the number of people with access. From that figure follows the serving approach, then the hardware, and only then the model.

For trial setups and the question of basic suitability we use Ollama ourselves, because it is the fastest tool for that. What goes to production runs on vLLM.

What we do not do is scale up an Ollama setup as user numbers grow. The switch belongs made before waiting times become noticeable, not after.

Further guides

Unclear how many people your planned hardware supports? We work it through against model, context length and concurrency before anything is purchased. Book a call

Sources

Enquiry

Size local AI for multiple users

We size hardware, model and inference server against your actual concurrency and take on operations.

What is the situation?

How should we get back to you?

Frequently Asked Questions

Answers to important questions about this topic

They sit at different levels. llama.cpp is an inference engine, Ollama a usability layer on top of it that uses llama.cpp, and vLLM a standalone serving system for GPU operation. Comparing Ollama against llama.cpp therefore compares a wrapper with its contents.

Because it is built for concurrent requests and batches them rather than processing them one after another. Published comparisons cite factors of ten and more under concurrency against the other two. With a single user that advantage largely disappears.

For small groups with occasional use it works. As soon as several people work at the same time, waiting becomes noticeable because requests are served sequentially. Moving to vLLM is then not fine-tuning but the difference between usable and unusable.

vLLM is under the Apache License 2.0, Ollama and llama.cpp under MIT. All three are unrestricted open source with no separated proprietary components.

In practice yes. vLLM is designed for GPU serving, and its advantage comes from techniques that use GPU memory efficiently. For CPU-only operation, llama.cpp is the right choice.

Used directly yes, by a low double-digit percentage, because Ollama as a wrapper adds overhead. For a single user that difference is barely noticeable in daily use and rarely outweighs Ollama's easier handling.

Ollama, without hesitation. Installation and a first model take minutes. The mistake is not starting there but going to production without having checked the user count.

No, and most teams do not. Ollama for development and demos plus vLLM for production is a common and sensible split.

Little, if the application talks to an OpenAI-compatible API, because both offer one. Check model names, quantisation format differences and behaviour with long contexts.

Like vLLM, SGLang is a serving system for many concurrent requests, licensed under Apache 2.0 and hosted by the non-profit organisation LMSYS. Its core technique, RadixAttention, keeps shared prefixes such as system prompts or conversation histories in the cache and reuses them across requests. As of October 2026 the current version is v0.5.20 from 18 September 2026.

Not across the board. The 2023 SGLang paper reports up to 6.4 times the throughput of the systems available at the time, measured on workloads with many shared prefixes. vLLM has since gained automatic prefix caching as well. Which system is faster depends on model, hardware, version and load profile and can only be measured with your own load.

Largely yes. According to its documentation SGLang supports NVIDIA GPUs including RTX PRO 6000 and DGX Spark, AMD Instinct, Google TPU, Intel GPUs and CPUs, and Apple Silicon. Since v0.5.20 there are no more images and wheels for CUDA 12; v0.5.19 was the last version for it.

vLLM, together with Open WebUI as the interface. That follows from the purpose: a device for a team has to serve several concurrent requests, and that is exactly what vLLM is built for.

Timo Wevelsiep

Written by

Timo Wevelsiep

Co-Founder & CEO

Co-Founder of WZ-IT. Specialized in cloud infrastructure, open-source platforms and managed services for SMEs and enterprise clients worldwide.

LinkedIn

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back by the next business day at the latest.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.