[email protected]
Interactive simulator

LLM Token/s Simulator: Qwen3.6-35B-A3B NVFP4 on GB10

Is one AI Cube enough for your team, and how does local AI feel in daily use? Experience measured output speed for nvidia/Qwen3.6-35B-A3B-NVFP4 in a single chat and under concurrent load on a GB10-based AI Cube.

Cloud Wolke HerausforderungServer NachhaltigkeitIT Beratung Service Consulting SoforthilfeTimo Wevelsiep Robin ZinsExperten für Innovation Migration AWSHetzner Hosting zuverlässig

Simulator

Try tokens per second yourself

Set speed and answer length, start the simulation and read along. The presets match measured values.

108 tok/s
10120 tok/s

Qwen presets show measured per-chat speed in a public GB10 test dated 19 July 2026. The workload used 1,024 input and 256 output tokens. We validate model, context and concurrency separately for your deployment.

500 Token
503.000 Token
Output0 / 500 Token
Start the simulation to read the answer at the chosen speed.
Estimated
4.63 s
Elapsed
0.00 s
Current
0.0 tok/s

Current measurements

Qwen3.6-35B-A3B-NVFP4 on one AI Cube

Visible speed per chat decreases under concurrent load while total system throughput increases. Both matter when sizing the system.

AI Cube · 1 × NVIDIA GB10 · 128 GB Unified Memory

nvidia/Qwen3.6-35B-A3B-NVFP4

  • ≈ 108 tok/s · Single chat
  • ≈ 43 tok/s · 8 concurrent chats · per chat
  • ≈ 320 tok/s · 8 concurrent chats · aggregate
  • Architecture: 35B total · 3B active
  • Checkpoint: NVIDIA NVFP4
  • Runtime: vLLM · FP8 KV-Cache · MTP-2
  • Test workload: 1.024 Input · 256 Output
Alternative reference: GPT-OSS 20B reached around 91 tok/s in a single chat. Runtime and quantisation differ, so this is not a direct quality or model comparison.
Output speed under concurrent load
Concurrent chatsPer chatAggregateTTFT p50
1≈ 108 tok/s≈ 100 tok/s178 ms
2≈ 85 tok/s≈ 158 tok/s209 ms
4≈ 50 tok/s≈ 189 tok/s245 ms
8≈ 43 tok/s≈ 320 tok/s282 ms
16≈ 29 tok/s≈ 432 tok/s447 ms

Use cases

Typical use cases: how long an answer takes

From around 30 tokens per second per chat, an answer is easy to read as it streams. How long it takes overall depends on its length.

  • AI chat

    Visible speed for each active conversation
    • Tokens: 50-250
    • Rec. Speed: 30+ tok/s
    • Duration: 0.5-8s
  • E-Mail

    Response time also depends on the requested length
    • Tokens: 200-500
    • Rec. Speed: 30+ tok/s
    • Duration: 2-17s
  • Report/Article

    Aggregate throughput matters most for long outputs
    • Tokens: 1000-3000
    • Rec. Speed: 30+ tok/s
    • Duration: 9-100s

Planning figure for one AI Cube: up to ten answers generated at the same time. It sits between the measured 8- and 16-chat points; we do not derive a falsely precise ten-chat figure from it. For teams of five to twenty people one Cube is usually enough; sizing uses the target model, context and knowledge sources.

Next step

AI inference on your own network: the AI Cube

The AI Cube brings the measured performance into your network, without cloud dependency and without external token costs. Once for the device and its setup, monthly for maintenance.

One-off

The AI Cube

from €6,490 excl. VATone-off, 1 TB version
  • NVIDIA GB10 with 128 GB unified memory, preconfigured for your network
  • Tuned language models and Open WebUI as the working interface
  • Delivered in ten working days, briefing included

Ongoing maintenance

AI Cube Care

from €349.90 excl. VAT / monththe first twelve months with the device, cancellable monthly after that
  • Updates for system, inference engine and models in the maintenance window
  • Monitoring and alerting around the clock
  • Ticket support with an answer by the next working day

More concurrent load than one Cube carries, or no room on your premises: managed GPU servers in a German data centre, from €699 excl. VAT / month. View managed GPU servers

All prices are net and exclude statutory VAT. The offers are addressed to businesses.

Token Speed FAQ

Tokens, influencing factors, practice and the AI Cube

Basics

Influencing Factors

Practical Application

Hardware & AI Cube

Reviews & projects

What clients say about working with us

From local AI integration to architecture, data sovereignty and ongoing operations.

WZ-IT moved our studio infrastructure from decentralised individual devices to a central platform: every site is securely connected via VPN, new devices are onboarded automatically and an entire site is provisioned from a template, without manual steps on location. What impressed me most is the breadth and depth of their knowledge: Timo and Robin are not a typical IT provider who sets up a server and leaves. The two of them think their way into highly complex infrastructure and software topics, work through every requirement we put in front of them, and build networking, provisioning and operations so that everything fits together in the end. WZ-IT is an excellent partner for complex software, network and architecture projects.
Steve KirchnerManaging Director, nextGYM GmbH
View project

International

Built in Germany's Ruhr Valley. Running worldwide.

WZ-IT designs, develops and operates infrastructure and software for clients in Germany and internationally. We deliver projects remotely and continue supporting them in ongoing operations after go-live.

Selected projects

Read client reviews

  • Integrate a local AI solution
  • Modernize your infrastructure - sovereign
  • Design an open-source AI architecture
  • Plan a sovereign open-source stack
  • Secure your Proxmox & backup setup
  • Modernize your legacy software
  • Cut cloud cost - up to −81%
  • Build a high-availability Proxmox cluster
  • Virtualize with Managed Proxmox
  • Get collaboration fully managed
  • Ship your prototype to production
  • Connect sites and clusters securely

Enquiry

Turn the result into a defensible configuration

The simulator shows reference values from measurements. For a proposal, we review the model, quantisation, workload, and integrations together.

  • Straight with Timo and Robin - no sales team, no pitch
  • An honest take, including when we are not the right fit
  • Concrete next steps for infrastructure, software or AI

No risk: worst case, you leave with a clearer understanding of your project than before.

Timo and Robin, founders of WZ-IT

Which infrastructure should be assessed?

Copy the main assumptions from the calculator into the text field.

We usually respond within one business day. Please do not send access credentials yet.

WZ-IT's advice on our Azure migration was technically sound and completely non-binding right from the intro call - we took away a great deal.
Jakob ÖschlbergerInno7 GmbH