Simulate token generation speed for Large Language Models and understand how different speeds affect user experience.
The AI Cube presets are conservative guidance values based on publicly available GB10/DGX Spark benchmarks - not a measurement of your workload. Binding figures come from a benchmark with your target model, context length and concurrency.
0.00s
0 tok/s
Local AI inference without cloud dependency. Full data control.
ASUS/NVIDIA GB10 Appliance
| Model | WZ-IT AI Cube |
|---|---|
GPT-OSS 20B ~20 Billion Parameters | 80-90 tok/s |
GPT-OSS 120B ~120 Billion Parameters | 35-60 tok/s |
* Conservative guidance based on publicly available GB10/DGX Spark benchmarks. Performance depends on model, context length, backend and concurrency.
Quick responses for fluid conversation
Sufficient for email generation
Longer texts with good performance
Everything you need to know about token generation
Topics
Token generation speed (measured in tokens per second or tok/s) indicates how fast an AI model can generate text. One token corresponds to approximately 4 characters or 0.75 words. At 100 tok/s, about 75 words are generated per second.
Simulation helps developers and businesses understand how different speeds affect user experience. Slow generation (under 30 tok/s) feels sluggish, while fast generation (over 100 tok/s) provides a fluid experience.
A token is the smallest unit an LLM processes. It can be a word, word part, or punctuation mark. 'Hello' is one token, 'championship' might be split into 'champion' + 'ship'. Most LLMs use about 1 token per 4 characters.
Key factors include: 1) GPU performance and VRAM, 2) Model size (7B, 70B, 120B parameters), 3) Quantization (FP16, INT8, INT4), 4) Batch size, 5) Context length, 6) Inference backend (vLLM, Ollama, TensorRT).
Larger models are slower. A 7B model can achieve 200+ tok/s, a 70B model about 50-100 tok/s, and a 120B model typically 30-60 tok/s on consumer hardware. However, response quality increases with model size.
No, speed varies. At the beginning (prefill phase), generation is often slower, then stabilizes. Query complexity, context length, and system load also affect speed.
Under 20 tok/s: Noticeably slow, frustrating. 20-50 tok/s: Acceptable for most applications. 50-100 tok/s: Good experience, feels fluid. Over 100 tok/s: Excellent, text appears almost instantly.
Chatbots: 50-100 tok/s for fluid conversation. Document generation: 30-50 tok/s sufficient. Real-time translation: 100+ tok/s recommended. Code completion: 100+ tok/s for best developer experience.
1) Use better GPU with more VRAM, 2) Quantization (INT4/INT8) for smaller models, 3) Use optimized inference engines like vLLM, 4) Batch processing for multiple requests, 5) Optimize KV-cache, 6) Enable continuous batching.
WZ-IT AI Cube based on GB10: conservative guidance is around 80-90 tok/s with GPT-OSS 20B and around 35-60 tok/s with GPT-OSS 120B. Values depend on context length, quantization, backend and concurrency.
Cloud APIs have network latency, rate limits, and share resources with other users. Local hardware like the AI Cube offers dedicated resources, no network delay, and consistent performance without queues.
Cloud APIs charge per token; local hardware has a one-off purchase cost and no token costs afterwards. The volume at which that flips depends on the model, the provider price and your utilisation - we work it through with your numbers in the intro call.
The AI Cube offers local AI inference on owned hardware - without cloud dependency and external token costs.
ml&s speaks about integrating a local AI solution. The other voices cover architecture, data sovereignty and operations - exactly the maturity an AI project needs to reach production.
Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.
Timo Wevelsiep & Robin Zins
Managing Directors of WZ-IT
