WZ-IT Logo

Multi-GPU inference without NVLink: tensor parallelism, pipeline parallelism or replicas

Timo WevelsiepTimo Wevelsiep•Updated: 30.09.2026

Editorial note: Versions, commands and prices may change. Please verify critical steps independently before production use. This guide does not replace individual consulting.

Larger model or more throughput than one GPU can deliver? WZ-IT runs managed GPU servers with NVIDIA RTX PRO 4000 Blackwell (24 GB) and RTX PRO 6000 Blackwell Max-Q (96 GB), including vLLM, Open WebUI and monitoring. We plan multi-GPU configurations to match model and load. View GPU servers · Book a call

A server with two GPUs does not automatically deliver twice the performance. What matters is how the inference software distributes the model across the cards and how much bandwidth is available between the cards. Data centre GPUs such as the H100 or B200 are coupled via NVLink. The RTX PRO Blackwell cards found in many business servers communicate over PCIe only. That shifts the trade-off between tensor parallelism, pipeline parallelism and several independent instances (replicas). This article explains the three approaches, the corresponding vLLM parameters and a decision rule for systems without NVLink. As of October 2026.

Table of contents

NVLink is NVIDIA's direct GPU-to-GPU interconnect. According to NVIDIA, the fifth generation (Blackwell) reaches 1,800 GB/s per GPU and the fourth generation (Hopper) 900 GB/s (NVIDIA, NVLink). For comparison, NVIDIA lists 128 GB/s for the PCIe Gen5 connection of an H100 (NVIDIA, H100). That figure covers both directions of an x16 link combined.

The RTX PRO Blackwell cards in the following table have no NVLink connector. Their datasheets list PCIe as the only system interface:

GPU Memory Memory bandwidth System interface NVLink Source
RTX PRO 6000 Blackwell Workstation Edition 96 GB GDDR7 ECC 1,792 GB/s PCIe 5.0 x16 not listed Datasheet
RTX PRO 6000 Blackwell Max-Q 96 GB GDDR7 ECC 1,792 GB/s PCIe 5.0 x16 not listed Datasheet
RTX PRO 4000 Blackwell 24 GB GDDR7 ECC 672 GB/s PCIe 5.0 x16 not listed Datasheet
RTX PRO 4000 Blackwell SFF Edition 24 GB GDDR7 ECC 432 GB/s PCIe 5.0 x8 not listed Datasheet

The order of magnitude matters: inside an RTX PRO 6000, memory moves 1,792 GB/s, while PCIe 5.0 x16 between two cards provides a calculated 60-plus GB/s per direction. Any data exchange between the GPUs is therefore more than an order of magnitude slower than an access to local memory. Whether the cards in a server are actually connected with full lane width and without a detour through the second CPU depends on the mainboard (see Check the topology).

Three ways to put one model on several GPUs

vLLM offers three basic approaches that can be combined (vLLM, Parallelism and Scaling):

Approach What is split Communication between GPUs Benefit vLLM parameter
Tensor parallelism (TP) Every layer is sliced across all GPUs In every layer (all-reduce) Large model fits, a single request gets faster --tensor-parallel-size (-tp)
Pipeline parallelism (PP) Layers are distributed across GPUs in blocks Only at block boundaries Large model fits, less traffic than TP --pipeline-parallel-size (-pp)
Data parallelism (DP) / replicas Nothing, each GPU holds a full copy None (for dense models) More concurrent requests, fault tolerance --data-parallel-size (-dp) or separate instances
Expert parallelism (EP) Experts of an MoE model are distributed across GPUs Token routing to the experts Distribute MoE models --enable-expert-parallel (-ep)

All of these parameters default to 1 or are off by default in vLLM (vLLM, Engine Arguments). The number of GPUs required is the product: -dp 2 -tp 2 needs four GPUs.

Tensor parallelism is the approach most guides mention first. The vLLM documentation recommends it when a model does not fit on one GPU but fits on one server.

Pipeline parallelism places, for example, the first half of the layers on GPU 0 and the second half on GPU 1. Only the intermediate results at the handover point travel between the cards. A single request does not get faster, because the GPUs work one after the other. With many concurrent requests, both stages work in parallel on different requests.

Replicas do not split the model; they are several complete instances. vLLM provides --data-parallel-size for this, with a shared API endpoint and internal load balancing; it works for dense and MoE models (vLLM, Data Parallel Deployment). Alternatively, two separate vLLM processes run behind a gateway such as LiteLLM.

Expert parallelism only concerns mixture-of-experts models. The experts are distributed across the GPUs instead of slicing every expert with TP; the EP size is TP times DP (vLLM, Expert Parallel Deployment).

Tensor parallelism goes back to the method from Megatron-LM: the weight matrices of attention and MLP are split by columns or rows, each GPU computes its part, and after attention and after the MLP the partial results are combined with an all-reduce (Shoeybi et al., Megatron-LM). That makes two synchronisation points per transformer layer. A model with 60 layers therefore synchronises around 120 times for every generated token.

During token generation (decode), the amount of data per step is small, but every synchronisation carries a fixed latency. Processing long inputs (prefill) moves large amounts of data. Over NVLink, both effects barely matter. Over PCIe, the GPU measurably waits for the other card, and the share of waiting time grows with every additional TP GPU.

The vLLM documentation draws an explicit conclusion: if the GPUs in a server lack an NVLink interconnect, for example the L40S, pipeline parallelism should be used instead of tensor parallelism for higher throughput and lower communication overhead (vLLM, Parallelism and Scaling). The RTX PRO Blackwell cards fall into the same category.

This does not make TP over PCIe unusable. With two GPUs the overhead remains manageable, and TP shortens the response time of a single request, which PP does not. The gap widens with TP=4 and TP=8 and with the number of concurrent requests.

The decision rule: split as little as possible

For servers without NVLink the choice comes down to one rule: distribute the model across as few GPUs as needed for it to fit with enough KV cache, and use the remaining GPUs for additional replicas. How much memory weights and KV cache need is worked out in GPU and VRAM sizing for LLMs; which format shrinks the weights is explained in LLM quantization.

Starting point Recommendation for 2 GPUs without NVLink Reason
Model plus KV cache fits on one GPU Two replicas (-dp 2 or two instances) No traffic between GPUs, throughput scales, failure of one GPU is tolerable
Model fits on one GPU, but too little KV cache for context or user count Check quantization or FP8 KV cache first, otherwise TP=2 Replicas with too small a KV cache lose their advantage
Model does not fit on one GPU TP=2 or PP=2, measure both TP for shorter response times, PP for more throughput under many requests
Short response time per request takes priority TP=2 Only TP reduces compute time per token
Mixture-of-experts model across two GPUs Measure TP=2 with and without --enable-expert-parallel The effect depends on model and load
Several small models on one large GPU MIG partitions (RTX PRO 6000: up to 4 x 24 GB) Replicas or different models isolated on one card

An example of the third case: Qwen3.5-122B-A10B in FP8 occupies considerably more than the 96 GB of an RTX PRO 6000 with its weights alone. Two cards with TP=2 hold weights and KV cache together; the setup is documented in vLLM with Qwen3.5-122B on two RTX PRO 6000.

An example of the first case: a model with around 30 billion parameters in FP8 or 4-bit fits on a 96 GB card with plenty of room for KV cache. With two such GPUs, two instances deliver more throughput than one instance with TP=2.

vLLM configuration for two GPUs

The following commands show the three variants with vllm serve. Model name, context length and memory fraction are placeholders that have to match the model.

Tensor parallelism: one model across both GPUs

vllm serve <model> \
  --tensor-parallel-size 2 \
  --max-model-len 32768 \
  --gpu-memory-utilization 0.90

Pipeline parallelism: layers distributed in blocks

vllm serve <model> \
  --pipeline-parallel-size 2 \
  --max-model-len 32768

Replicas via data parallelism: two copies behind one endpoint

vllm serve <model> \
  --data-parallel-size 2 \
  --max-model-len 32768

Replicas as separate instances

CUDA_VISIBLE_DEVICES=0 vllm serve <model> --port 8000
CUDA_VISIBLE_DEVICES=1 vllm serve <model> --port 8001

Separate instances can be restarted and updated individually. A gateway with load balancing and health checks distributes the requests, for example LiteLLM. With --data-parallel-size, vLLM handles the distribution itself; queue limits then apply to the whole server, not per rank (vLLM, Data Parallel Deployment).

Further parameters that matter on PCIe systems:

Parameter Effect When relevant
--enable-expert-parallel Distribute the experts of an MoE model instead of slicing them with TP MoE models with TP or DP greater than 1
--disable-custom-all-reduce Disable the custom all-reduce kernel and fall back to NCCL Troubleshooting hangs or startup failures with TP
--distributed-executor-backend mp (multiprocessing) or ray One server: mp; several nodes: depends on the setup
--max-num-seqs Maximum number of sequences processed concurrently Balance throughput against KV cache demand

The descriptions come from the engine arguments reference (vLLM, Engine Arguments). How vLLM differs from Ollama overall is shown in vLLM vs. Ollama.

What a public measurement shows

Reliable published measurements of TP versus replicas on PCIe systems are rare. One reproducible series is vllm-topology-bench. It compares one instance with TP=4 against two replicas with TP=2 each on four RTX 3090 GPUs without NVLink:

Item Value
Hardware 4 x RTX 3090 (24 GB), no NVLink
Model Qwen3.6-35B-A3B, AWQ-quantised (weights around 23 GB)
vLLM version 0.21.0
Context length 8,192 tokens, 512 tokens input, 256 tokens output
Throughput at 1 request 111 (TP=4) vs. 116 tokens/s (2 x TP=2)
Throughput at 128 concurrent requests 610 (TP=4) vs. 1,437 tokens/s (2 x TP=2)
Time to first token at 16 concurrent requests 3.9 s (TP=4) vs. 1.4 s (2 x TP=2)

The pattern confirms the rule: with a single request both variants are almost level; as load increases, the variant with more TP GPUs saturates earlier. One copy per card (4 x TP=1) was not possible in this measurement because the weights almost fill the 24 GB cards.

The project names its own limitations: one run per data point, consumer cards, short context length, a single MoE model. The absolute numbers do not transfer to RTX PRO Blackwell or other models. The direction does, and it should be checked on your own system with your own load before committing to a layout.

Check and measure the topology

How well two GPUs communicate depends not only on the card but also on the server. Three checks belong in front of every multi-GPU configuration:

  1. Show the connection between the GPUs. nvidia-smi topo -m displays the topology matrix. If the GPUs sit on the same PCIe switch or root complex, the path is short. If the matrix shows SYS between them, traffic crosses the link between two CPUs, which slows TP down further.
  2. Check lane width and generation. nvidia-smi -q shows the current PCIe generation and link width. A card connected with only x8 or PCIe 4.0 halves the available bandwidth.
  3. Measure with realistic load. Compare throughput and time to first token at the expected number of concurrent users and the actual input length, for TP=2, PP=2 and replicas where the model allows it. vLLM ships its own benchmark tool for this (vllm bench serve).

When starting with TP, NCCL logs which path the GPUs use to communicate. The vLLM documentation describes how to read this output (vLLM, Parallelism and Scaling). Connecting several servers instead of several GPUs in one server raises the same questions at network level; for GB10 systems this is covered in Connect AI Cubes with ConnectX-7.

What this means for your project

Two GPUs without NVLink are a sound basis for production use if the split matches the goal. For more users with a model that fits on one card, replicas are the direct route. For models that exceed one card, TP=2 is the common standard and PP=2 the alternative that can pay off under high load. TP across four or more PCIe GPUs should only be chosen after measuring it yourself.

WZ-IT provides managed GPU servers with a dedicated NVIDIA RTX PRO 4000 Blackwell (24 GB, from €699 net per month) or RTX PRO 6000 Blackwell Max-Q (96 GB, from €1,799 net per month), each with vLLM, Open WebUI and monitoring. We plan multi-GPU configurations, separate instances behind a gateway and redundant model endpoints to match model, context length and user count, and itemise them in the quote. For on-premises operation, the AI Cube is the alternative. Support, consulting and implementation by WZ-IT.

Which model is a candidate in the first place is covered in Which LLM to self-host?. How many users a server can carry is described in Size a local AI server by user count, and the differences between inference engines are compared in vLLM, Ollama and llama.cpp compared.

Rather have it operated?

You'd rather not run Local AI for Business yourself? WZ-IT handles setup, operations and maintenance - privacy-focused from Germany.

Enquiry

Assess local AI for your use case

Start with the AI Cube or have us assess a custom AI platform, knowledge connection, or integration.

How should we get back to you?

Frequently Asked Questions

Answers to the most important questions

No. NVIDIA's datasheets for the RTX PRO 6000 Blackwell Workstation Edition and the Max-Q Workstation Edition list PCIe 5.0 x16 as the system interface and do not list an NVLink connector. The same applies to the RTX PRO 4000 Blackwell (PCIe 5.0 x16) and its SFF Edition (PCIe 5.0 x8). Several of these cards in one server exchange data over PCIe (as of October 2026).

No. Each GPU keeps its own memory. A model that needs more than 96 GB can only be used if the inference software splits it, for example with tensor parallelism or pipeline parallelism in vLLM. The combined memory is then available for weights and KV cache, but every split creates traffic between the cards.

No. Tensor parallelism splits every layer across both GPUs and synchronises the partial results in every layer with an all-reduce operation. Without NVLink this synchronisation runs over PCIe and costs time. Compute doubles, but total throughput rises considerably less than twofold. If the model fits on one GPU, two independent instances usually deliver more throughput.

No. vLLM supports tensor parallelism, pipeline parallelism and data parallelism over PCIe as well. For GPUs without NVLink, the vLLM documentation even recommends considering pipeline parallelism instead of tensor parallelism because the communication overhead is lower. NVLink improves scaling but is not a requirement.

When the model, including the KV cache for the required context length, fits on a single GPU. Each GPU then handles its own requests, there is no synchronisation between the cards and throughput scales with the number of instances. If one GPU fails, the other keeps answering. Tensor parallelism is the choice when the model would not fit otherwise or when a single request has to be answered faster.

--tensor-parallel-size 2 (short -tp 2) splits every layer across two GPUs. --pipeline-parallel-size 2 (-pp 2) distributes the layers in blocks. --data-parallel-size 2 (-dp 2) starts two complete copies behind one endpoint. For mixture-of-experts models, --enable-expert-parallel distributes the experts instead of slicing them with tensor parallelism. All parameters default to 1.

It measures an AWQ-quantised Qwen3.6-35B-A3B with vLLM 0.21.0 on four RTX 3090 GPUs without NVLink. Two replicas with TP=2 each reached 1,437 tokens per second at 128 concurrent requests, a single instance with TP=4 only 610. It is a single measurement series on consumer cards with an 8,192-token context. The direction carries over, the numbers do not apply to other GPUs and models.

Yes. According to NVIDIA's datasheet, the RTX PRO 6000 Blackwell Max-Q supports Multi-Instance GPU (MIG) with up to four instances of 24 GB each or two of 48 GB each. Each instance behaves like a separate GPU. This is the counterpart to multi-GPU: several replicas on one card instead of one model on several cards.

More on Local AI for Business

Contact

Let's Talk About Your Idea

Whether a specific IT challenge or just an idea - we look forward to the exchange. In a brief conversation, we'll evaluate together if and how your project fits with WZ-IT.

Arrange a callback

Callback

Arrange a callback

Leave your number and we will call back — at the latest on the next business day.

For a longer conversation you can book an appointment instead.

Companies worldwide trust WZ-IT

  • ml&s
  • Rekorder
  • Keymate
  • Führerscheinmacher
  • SolidProof
  • ARGE
  • Boese VA
  • nextGYM
  • SweetConnect GmbH
  • Golem.de
  • Millenium
  • Paritel
  • Yonju
  • EVADXB
  • Mr. Clipart
  • Aphy AG
  • Negosh
  • ABCO Water Systems
1/3 - Topic Selection33%

What is your inquiry about?

First select the service area that best matches your project.