AI Cube · 1 × NVIDIA GB10 · 128 GB Unified Memory
nvidia/Qwen3.6-35B-A3B-NVFP4
- ≈ 108 tok/s · Single chat
- ≈ 43 tok/s · 8 concurrent chats · per chat
- ≈ 320 tok/s · 8 concurrent chats · aggregate
- Architecture: 35B total · 3B active
- Checkpoint: NVIDIA NVFP4
- Runtime: vLLM · FP8 KV-Cache · MTP-2
- Test workload: 1.024 Input · 256 Output



