How many users can a local AI server support?
Timo Wevelsiep•Updated: 15.08.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Does the AI Cube Pro fit your user load? We align the model and expected concurrency before delivery. If measured load exceeds the platform, we add AI Cubes or size a GPU server. Configure the AI Cube Pro
A local AI server has no credible fixed limit such as “50 users”. One hundred accounts may be easy when only a few people ask short questions at once. Ten concurrent users can be demanding when they process large documents and expect long answers. Capacity planning therefore starts with concurrency rather than account count.
Short answer: which setup fits which usage profile?
| Usage profile | Sensible first step |
|---|---|
| individual or sporadically active users with short chats | one AI Cube with an agreed model and documented baseline test |
| a department running document searches concurrently on a regular basis | load-test one AI Cube, define capacity headroom, and add a second model endpoint when required |
| many parallel API calls, long contexts, or several active models | size AI Cube Custom or a dedicated GPU server from the workload |
This table is not a fixed performance commitment. It identifies the architecture to test first. A defensible user count only follows from the exact model, representative tasks, and acceptable waiting time.
Five factors that determine throughput
- Model and quantisation: larger weights use more memory and compute per token. The guide to LLMs on 128 GB of unified memory maps the relevant model classes.
- Input context: long chats and document extracts increase prefill work and KV-cache use.
- Response length: a short answer occupies capacity for far less time than a long report.
- Concurrent requests: inference servers batch work, but every active sequence consumes resources.
- Expected latency: a team may accept a minute for analysis but only a few seconds to first output in chat.
Three user counts instead of one
Record registered users, daily active users, and users active concurrently during a peak. The third number drives hardware sizing. Also separate chat, document search, transcription, and automated API jobs, as they create different load profiles.
What we measure on the AI Cube Pro
The AI Cube Pro provides 128 GB of unified memory and a local platform prepared for multi-user operation. We agree a model for the use case before delivery. Capacity is then measured with representative requests:
- time to first token;
- visible output rate per request;
- aggregate throughput;
- memory use at typical and maximum context;
- queues, cancellations, and errors under peak load.
Results apply to the documented model release, quantisation, runtime, and test profile. Moving to a larger model can substantially alter capacity.
When a second AI Cube makes sense
For more concurrent users, two AI Cubes can process independent requests. A load balancer sends new work to available capacity. This differs from distributing one large model across two systems. Distributed inference uses ConnectX-7 and adds network and software configuration to the performance profile.
A second node does not automatically provide high availability. Model endpoints, interface, database, knowledge store, identity, and load balancer must also be redundant or recoverable.
When AI Cube Custom or a GPU server follows
A larger deployment is appropriate for consistently larger models, many long concurrent contexts, several resident models, or defined availability. We then size AI Cube Custom or GPU servers from the measured profile.
A practical load-test profile
- select five to ten typical tasks;
- define realistic prompt, document, and output sizes;
- test one, two, four, and eight concurrent workflows;
- assess quality and speed together;
- add peak periods and automated jobs;
- define explicit capacity headroom for growth and updates.
This turns a user count into an auditable capacity decision rather than a marketing number.
Sources and further reading
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube Pro or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
The number of registered accounts is not the bottleneck. Concurrent users, model, context, response length, and desired waiting time determine capacity. WZ-IT tests the intended workload instead of promising an unsupported fixed user limit.
They affect different resources. Model weights occupy a base amount of memory; context and concurrency expand the KV cache and determine throughput and latency.
For independent requests, a load balancer can distribute work across two systems and substantially increase aggregate throughput. A single model distributed across both systems has different performance and networking constraints.
Use representative prompts, documents, contexts, and concurrent requests. Measure time to first token, output rate, aggregate throughput, memory use, and error rate.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for confidentiality professions
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





