Deployed worldwide
WZ-IT Logo

How many users can a local AI server support?

Timo WevelsiepTimo WevelsiepUpdated: 15.08.2026

Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.

Does the AI Cube Pro fit your user load? We align the model and expected concurrency before delivery. If measured load exceeds the platform, we add AI Cubes or size a GPU server. Configure the AI Cube Pro

A local AI server has no credible fixed limit such as “50 users”. One hundred accounts may be easy when only a few people ask short questions at once. Ten concurrent users can be demanding when they process large documents and expect long answers. Capacity planning therefore starts with concurrency rather than account count.

Short answer: which setup fits which usage profile?

Usage profile Sensible first step
individual or sporadically active users with short chats one AI Cube with an agreed model and documented baseline test
a department running document searches concurrently on a regular basis load-test one AI Cube, define capacity headroom, and add a second model endpoint when required
many parallel API calls, long contexts, or several active models size AI Cube Custom or a dedicated GPU server from the workload

This table is not a fixed performance commitment. It identifies the architecture to test first. A defensible user count only follows from the exact model, representative tasks, and acceptable waiting time.

Five factors that determine throughput

  1. Model and quantisation: larger weights use more memory and compute per token. The guide to LLMs on 128 GB of unified memory maps the relevant model classes.
  2. Input context: long chats and document extracts increase prefill work and KV-cache use.
  3. Response length: a short answer occupies capacity for far less time than a long report.
  4. Concurrent requests: inference servers batch work, but every active sequence consumes resources.
  5. Expected latency: a team may accept a minute for analysis but only a few seconds to first output in chat.

Three user counts instead of one

Record registered users, daily active users, and users active concurrently during a peak. The third number drives hardware sizing. Also separate chat, document search, transcription, and automated API jobs, as they create different load profiles.

What we measure on the AI Cube Pro

The AI Cube Pro provides 128 GB of unified memory and a local platform prepared for multi-user operation. We agree a model for the use case before delivery. Capacity is then measured with representative requests:

  • time to first token;
  • visible output rate per request;
  • aggregate throughput;
  • memory use at typical and maximum context;
  • queues, cancellations, and errors under peak load.

Results apply to the documented model release, quantisation, runtime, and test profile. Moving to a larger model can substantially alter capacity.

When a second AI Cube makes sense

For more concurrent users, two AI Cubes can process independent requests. A load balancer sends new work to available capacity. This differs from distributing one large model across two systems. Distributed inference uses ConnectX-7 and adds network and software configuration to the performance profile.

A second node does not automatically provide high availability. Model endpoints, interface, database, knowledge store, identity, and load balancer must also be redundant or recoverable.

When AI Cube Custom or a GPU server follows

A larger deployment is appropriate for consistently larger models, many long concurrent contexts, several resident models, or defined availability. We then size AI Cube Custom or GPU servers from the measured profile.

A practical load-test profile

  1. select five to ten typical tasks;
  2. define realistic prompt, document, and output sizes;
  3. test one, two, four, and eight concurrent workflows;
  4. assess quality and speed together;
  5. add peak periods and automated jobs;
  6. define explicit capacity headroom for growth and updates.

This turns a user count into an auditable capacity decision rather than a marketing number.

Sources and further reading

Rather have it operated?

You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.

Enquiry

Assess local AI for your use case

Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.

How should we get back to you?

Frequently Asked Questions

Answers to the most important questions

The number of registered accounts is not the bottleneck. Concurrent users, model, context, response length, and desired waiting time determine capacity. WZ-IT tests the intended workload instead of promising an unsupported fixed user limit.

They affect different resources. Model weights occupy a base amount of memory; context and concurrency expand the KV cache and determine throughput and latency.

For independent requests, a load balancer can distribute work across two systems and substantially increase aggregate throughput. A single model distributed across both systems has different performance and networking constraints.

Use representative prompts, documents, contexts, and concurrent requests. Measure time to first token, output rate, aggregate throughput, memory use, and error rate.

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Email
[email protected]
Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • Maho Management
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/2 - Topic Selection50%

What is your inquiry about?

Select one or more areas where we can support you.