Multi-GPU inference without NVLink: tensor parallelism, pipeline parallelism or replicas
Timo Wevelsiep•Updated: 30.09.2026Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.
Larger model or more throughput than one GPU can deliver? WZ-IT runs managed GPU servers with NVIDIA RTX PRO 4000 Blackwell (24 GB) and RTX PRO 6000 Blackwell Max-Q (96 GB), including vLLM, Open WebUI and monitoring. We plan multi-GPU configurations to match model and load. View GPU servers · Book a call
A server with two GPUs does not automatically deliver twice the performance. What matters is how the inference software distributes the model across the cards and how much bandwidth is available between the cards. Data centre GPUs such as the H100 or B200 are coupled via NVLink. The RTX PRO Blackwell cards found in many business servers communicate over PCIe only. That shifts the trade-off between tensor parallelism, pipeline parallelism and several independent instances (replicas). This article explains the three approaches, the corresponding vLLM parameters and a decision rule for systems without NVLink. As of October 2026.
Table of contents
- NVLink and PCIe on RTX PRO Blackwell
- Three ways to put one model on several GPUs
- Why tensor parallelism gets expensive without NVLink
- The decision rule: split as little as possible
- vLLM configuration for two GPUs
- What a public measurement shows
- Check and measure the topology
NVLink and PCIe on RTX PRO Blackwell
NVLink is NVIDIA's direct GPU-to-GPU interconnect. According to NVIDIA, the fifth generation (Blackwell) reaches 1,800 GB/s per GPU and the fourth generation (Hopper) 900 GB/s (NVIDIA, NVLink). For comparison, NVIDIA lists 128 GB/s for the PCIe Gen5 connection of an H100 (NVIDIA, H100). That figure covers both directions of an x16 link combined.
The RTX PRO Blackwell cards in the following table have no NVLink connector. Their datasheets list PCIe as the only system interface:
| GPU | Memory | Memory bandwidth | System interface | NVLink | Source |
|---|---|---|---|---|---|
| RTX PRO 6000 Blackwell Workstation Edition | 96 GB GDDR7 ECC | 1,792 GB/s | PCIe 5.0 x16 | not listed | Datasheet |
| RTX PRO 6000 Blackwell Max-Q | 96 GB GDDR7 ECC | 1,792 GB/s | PCIe 5.0 x16 | not listed | Datasheet |
| RTX PRO 4000 Blackwell | 24 GB GDDR7 ECC | 672 GB/s | PCIe 5.0 x16 | not listed | Datasheet |
| RTX PRO 4000 Blackwell SFF Edition | 24 GB GDDR7 ECC | 432 GB/s | PCIe 5.0 x8 | not listed | Datasheet |
The order of magnitude matters: inside an RTX PRO 6000, memory moves 1,792 GB/s, while PCIe 5.0 x16 between two cards provides a calculated 60-plus GB/s per direction. Any data exchange between the GPUs is therefore more than an order of magnitude slower than an access to local memory. Whether the cards in a server are actually connected with full lane width and without a detour through the second CPU depends on the mainboard (see Check the topology).
Three ways to put one model on several GPUs
vLLM offers three basic approaches that can be combined (vLLM, Parallelism and Scaling):
| Approach | What is split | Communication between GPUs | Benefit | vLLM parameter |
|---|---|---|---|---|
| Tensor parallelism (TP) | Every layer is sliced across all GPUs | In every layer (all-reduce) | Large model fits, a single request gets faster | --tensor-parallel-size (-tp) |
| Pipeline parallelism (PP) | Layers are distributed across GPUs in blocks | Only at block boundaries | Large model fits, less traffic than TP | --pipeline-parallel-size (-pp) |
| Data parallelism (DP) / replicas | Nothing, each GPU holds a full copy | None (for dense models) | More concurrent requests, fault tolerance | --data-parallel-size (-dp) or separate instances |
| Expert parallelism (EP) | Experts of an MoE model are distributed across GPUs | Token routing to the experts | Distribute MoE models | --enable-expert-parallel (-ep) |
All of these parameters default to 1 or are off by default in vLLM (vLLM, Engine Arguments). The number of GPUs required is the product: -dp 2 -tp 2 needs four GPUs.
Tensor parallelism is the approach most guides mention first. The vLLM documentation recommends it when a model does not fit on one GPU but fits on one server.
Pipeline parallelism places, for example, the first half of the layers on GPU 0 and the second half on GPU 1. Only the intermediate results at the handover point travel between the cards. A single request does not get faster, because the GPUs work one after the other. With many concurrent requests, both stages work in parallel on different requests.
Replicas do not split the model; they are several complete instances. vLLM provides --data-parallel-size for this, with a shared API endpoint and internal load balancing; it works for dense and MoE models (vLLM, Data Parallel Deployment). Alternatively, two separate vLLM processes run behind a gateway such as LiteLLM.
Expert parallelism only concerns mixture-of-experts models. The experts are distributed across the GPUs instead of slicing every expert with TP; the EP size is TP times DP (vLLM, Expert Parallel Deployment).
Why tensor parallelism gets expensive without NVLink
Tensor parallelism goes back to the method from Megatron-LM: the weight matrices of attention and MLP are split by columns or rows, each GPU computes its part, and after attention and after the MLP the partial results are combined with an all-reduce (Shoeybi et al., Megatron-LM). That makes two synchronisation points per transformer layer. A model with 60 layers therefore synchronises around 120 times for every generated token.
During token generation (decode), the amount of data per step is small, but every synchronisation carries a fixed latency. Processing long inputs (prefill) moves large amounts of data. Over NVLink, both effects barely matter. Over PCIe, the GPU measurably waits for the other card, and the share of waiting time grows with every additional TP GPU.
The vLLM documentation draws an explicit conclusion: if the GPUs in a server lack an NVLink interconnect, for example the L40S, pipeline parallelism should be used instead of tensor parallelism for higher throughput and lower communication overhead (vLLM, Parallelism and Scaling). The RTX PRO Blackwell cards fall into the same category.
This does not make TP over PCIe unusable. With two GPUs the overhead remains manageable, and TP shortens the response time of a single request, which PP does not. The gap widens with TP=4 and TP=8 and with the number of concurrent requests.
The decision rule: split as little as possible
For servers without NVLink the choice comes down to one rule: distribute the model across as few GPUs as needed for it to fit with enough KV cache, and use the remaining GPUs for additional replicas. How much memory weights and KV cache need is worked out in GPU and VRAM sizing for LLMs; which format shrinks the weights is explained in LLM quantization.
| Starting point | Recommendation for 2 GPUs without NVLink | Reason |
|---|---|---|
| Model plus KV cache fits on one GPU | Two replicas (-dp 2 or two instances) |
No traffic between GPUs, throughput scales, failure of one GPU is tolerable |
| Model fits on one GPU, but too little KV cache for context or user count | Check quantization or FP8 KV cache first, otherwise TP=2 | Replicas with too small a KV cache lose their advantage |
| Model does not fit on one GPU | TP=2 or PP=2, measure both | TP for shorter response times, PP for more throughput under many requests |
| Short response time per request takes priority | TP=2 | Only TP reduces compute time per token |
| Mixture-of-experts model across two GPUs | Measure TP=2 with and without --enable-expert-parallel |
The effect depends on model and load |
| Several small models on one large GPU | MIG partitions (RTX PRO 6000: up to 4 x 24 GB) | Replicas or different models isolated on one card |
An example of the third case: Qwen3.5-122B-A10B in FP8 occupies considerably more than the 96 GB of an RTX PRO 6000 with its weights alone. Two cards with TP=2 hold weights and KV cache together; the setup is documented in vLLM with Qwen3.5-122B on two RTX PRO 6000.
An example of the first case: a model with around 30 billion parameters in FP8 or 4-bit fits on a 96 GB card with plenty of room for KV cache. With two such GPUs, two instances deliver more throughput than one instance with TP=2.
vLLM configuration for two GPUs
The following commands show the three variants with vllm serve. Model name, context length and memory fraction are placeholders that have to match the model.
Tensor parallelism: one model across both GPUs
vllm serve <model> \
--tensor-parallel-size 2 \
--max-model-len 32768 \
--gpu-memory-utilization 0.90
Pipeline parallelism: layers distributed in blocks
vllm serve <model> \
--pipeline-parallel-size 2 \
--max-model-len 32768
Replicas via data parallelism: two copies behind one endpoint
vllm serve <model> \
--data-parallel-size 2 \
--max-model-len 32768
Replicas as separate instances
CUDA_VISIBLE_DEVICES=0 vllm serve <model> --port 8000
CUDA_VISIBLE_DEVICES=1 vllm serve <model> --port 8001
Separate instances can be restarted and updated individually. A gateway with load balancing and health checks distributes the requests, for example LiteLLM. With --data-parallel-size, vLLM handles the distribution itself; queue limits then apply to the whole server, not per rank (vLLM, Data Parallel Deployment).
Further parameters that matter on PCIe systems:
| Parameter | Effect | When relevant |
|---|---|---|
--enable-expert-parallel |
Distribute the experts of an MoE model instead of slicing them with TP | MoE models with TP or DP greater than 1 |
--disable-custom-all-reduce |
Disable the custom all-reduce kernel and fall back to NCCL | Troubleshooting hangs or startup failures with TP |
--distributed-executor-backend |
mp (multiprocessing) or ray |
One server: mp; several nodes: depends on the setup |
--max-num-seqs |
Maximum number of sequences processed concurrently | Balance throughput against KV cache demand |
The descriptions come from the engine arguments reference (vLLM, Engine Arguments). How vLLM differs from Ollama overall is shown in vLLM vs. Ollama.
What a public measurement shows
Reliable published measurements of TP versus replicas on PCIe systems are rare. One reproducible series is vllm-topology-bench. It compares one instance with TP=4 against two replicas with TP=2 each on four RTX 3090 GPUs without NVLink:
| Item | Value |
|---|---|
| Hardware | 4 x RTX 3090 (24 GB), no NVLink |
| Model | Qwen3.6-35B-A3B, AWQ-quantised (weights around 23 GB) |
| vLLM version | 0.21.0 |
| Context length | 8,192 tokens, 512 tokens input, 256 tokens output |
| Throughput at 1 request | 111 (TP=4) vs. 116 tokens/s (2 x TP=2) |
| Throughput at 128 concurrent requests | 610 (TP=4) vs. 1,437 tokens/s (2 x TP=2) |
| Time to first token at 16 concurrent requests | 3.9 s (TP=4) vs. 1.4 s (2 x TP=2) |
The pattern confirms the rule: with a single request both variants are almost level; as load increases, the variant with more TP GPUs saturates earlier. One copy per card (4 x TP=1) was not possible in this measurement because the weights almost fill the 24 GB cards.
The project names its own limitations: one run per data point, consumer cards, short context length, a single MoE model. The absolute numbers do not transfer to RTX PRO Blackwell or other models. The direction does, and it should be checked on your own system with your own load before committing to a layout.
Check and measure the topology
How well two GPUs communicate depends not only on the card but also on the server. Three checks belong in front of every multi-GPU configuration:
- Show the connection between the GPUs.
nvidia-smi topo -mdisplays the topology matrix. If the GPUs sit on the same PCIe switch or root complex, the path is short. If the matrix showsSYSbetween them, traffic crosses the link between two CPUs, which slows TP down further. - Check lane width and generation.
nvidia-smi -qshows the current PCIe generation and link width. A card connected with only x8 or PCIe 4.0 halves the available bandwidth. - Measure with realistic load. Compare throughput and time to first token at the expected number of concurrent users and the actual input length, for TP=2, PP=2 and replicas where the model allows it. vLLM ships its own benchmark tool for this (
vllm bench serve).
When starting with TP, NCCL logs which path the GPUs use to communicate. The vLLM documentation describes how to read this output (vLLM, Parallelism and Scaling). Connecting several servers instead of several GPUs in one server raises the same questions at network level; for GB10 systems this is covered in Connect AI Cubes with ConnectX-7.
What this means for your project
Two GPUs without NVLink are a sound basis for production use if the split matches the goal. For more users with a model that fits on one card, replicas are the direct route. For models that exceed one card, TP=2 is the common standard and PP=2 the alternative that can pay off under high load. TP across four or more PCIe GPUs should only be chosen after measuring it yourself.
WZ-IT provides managed GPU servers with a dedicated NVIDIA RTX PRO 4000 Blackwell (24 GB, from €699 net per month) or RTX PRO 6000 Blackwell Max-Q (96 GB, from €1,799 net per month), each with vLLM, Open WebUI and monitoring. We plan multi-GPU configurations, separate instances behind a gateway and redundant model endpoints to match model, context length and user count, and itemise them in the quote. For on-premises operation, the AI Cube is the alternative. Support, consulting and implementation by WZ-IT.
Which model is a candidate in the first place is covered in Which LLM to self-host?. How many users a server can carry is described in Size a local AI server by user count, and the differences between inference engines are compared in vLLM, Ollama and llama.cpp compared.
Rather have it operated?
You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.
Enquiry
Assess local AI for your use case
Start with the AI Cube or have us assess a custom AI platform, knowledge connection, or integration.
Frequently Asked Questions
Answers to the most important questions
No. NVIDIA's datasheets for the RTX PRO 6000 Blackwell Workstation Edition and the Max-Q Workstation Edition list PCIe 5.0 x16 as the system interface and do not list an NVLink connector. The same applies to the RTX PRO 4000 Blackwell (PCIe 5.0 x16) and its SFF Edition (PCIe 5.0 x8). Several of these cards in one server exchange data over PCIe (as of October 2026).
No. Each GPU keeps its own memory. A model that needs more than 96 GB can only be used if the inference software splits it, for example with tensor parallelism or pipeline parallelism in vLLM. The combined memory is then available for weights and KV cache, but every split creates traffic between the cards.
No. Tensor parallelism splits every layer across both GPUs and synchronises the partial results in every layer with an all-reduce operation. Without NVLink this synchronisation runs over PCIe and costs time. Compute doubles, but total throughput rises considerably less than twofold. If the model fits on one GPU, two independent instances usually deliver more throughput.
No. vLLM supports tensor parallelism, pipeline parallelism and data parallelism over PCIe as well. For GPUs without NVLink, the vLLM documentation even recommends considering pipeline parallelism instead of tensor parallelism because the communication overhead is lower. NVLink improves scaling but is not a requirement.
When the model, including the KV cache for the required context length, fits on a single GPU. Each GPU then handles its own requests, there is no synchronisation between the cards and throughput scales with the number of instances. If one GPU fails, the other keeps answering. Tensor parallelism is the choice when the model would not fit otherwise or when a single request has to be answered faster.
--tensor-parallel-size 2 (short -tp 2) splits every layer across two GPUs. --pipeline-parallel-size 2 (-pp 2) distributes the layers in blocks. --data-parallel-size 2 (-dp 2) starts two complete copies behind one endpoint. For mixture-of-experts models, --enable-expert-parallel distributes the experts instead of slicing them with tensor parallelism. All parameters default to 1.
It measures an AWQ-quantised Qwen3.6-35B-A3B with vLLM 0.21.0 on four RTX 3090 GPUs without NVLink. Two replicas with TP=2 each reached 1,437 tokens per second at 128 concurrent requests, a single instance with TP=4 only 610. It is a single measurement series on consumer cards with an 8,192-token context. The direction carries over, the numbers do not apply to other GPUs and models.
Yes. According to NVIDIA's datasheet, the RTX PRO 6000 Blackwell Max-Q supports Multi-Instance GPU (MIG) with up to four instances of 24 GB each or two of 48 GB each. Each instance behaves like a separate GPU. This is the counterpart to multi-GPU: several replicas on one card instead of one model on several cards.
More on Local AI for Business
- The open-source LLM stack
- What is LiteLLM?
- What is Langfuse?
- What is vLLM?
- vLLM vs. Ollama
- What is RAG?
- Knowledge transfer during employee transitions
- Connect Open WebUI to Nextcloud (RAG with ACLs)
- What is local AI?
- Cloud AI vs. self-hosted
- Private ChatGPT for business
- AI sovereignty for companies
- Which LLM to self-host?
- Sizing GPU & VRAM
- Inference vs. Training
- Qdrant vs. pgvector
- The EU AI Act for companies
- Local AI for professional secrecy holders
- Processing documents with AI
- AI agents & automation
- RAG with permissions
- Chatbot or knowledge navigator?
- AI agents: permissions and approvals
- AI assistants and the works council
- GDPR-compliant AI: assessment criteria
- What does a local AI server cost?
- Buy or rent an AI server?
- Size a local AI server by users
- LLM models on 128 GB unified memory
- RAG with Nextcloud, SharePoint, and DMS
- Chunking for RAG
- Hybrid search and reranking
- Contextual retrieval
- Measuring RAG quality
- Open LLM licences for commercial use
- LLM quantization
- MCP in the enterprise
- Protection against prompt injection
- Multi-GPU inference without NVLink
- Text-to-SQL
- Provide secure remote access to local AI
- Connect AI Cubes with ConnectX-7
- Run Open WebUI as a production appliance
- Configure ASUS Ascent GX10 for business
- Configure NVIDIA DGX Spark for business
- Configure Acer Veriton GN100 for business
- Configure Dell Pro Max with GB10 for business
- Configure Gigabyte AI TOP ATOM for business
- Configure HP ZGX Nano G1n for business
- Configure Lenovo ThinkStation PGX for business
- Configure MSI EdgeXpert for business





