advanced 3 min answer

A training cluster has 512 NVIDIA H100-class GPUs in 64 eight-GPU nodes and the scheduler reports 96 GPUs free. Roughly how much collective bandwidth per GPU does a new 64-GPU job get, and what should the capacity report say instead of "96 GPUs free"?

nvidianvlinkncclfragmentationcapacity-modelling
Show the full answer Hide the answer

The assumptions, stated

Two published figures set the whole answer. In NVIDIA's HGX H100 generation (2023 datasheets), each GPU has roughly 900 GB/s of NVLink bandwidth to the other seven GPUs in its NVSwitch domain. Off the node, a DGX H100 carries eight ConnectX-7 adapters at up to 400 Gb/s each — about 50 GB/s per GPU. The ratio between inside the node and outside it is close to 18×.

Assume the 96 free GPUs are the ordinary result of a busy scheduler: two free on each of 48 nodes, because long-running jobs took six apiece.

The arithmetic

A 64-GPU job placed on two free GPUs per node spans 32 nodes. A ring all-reduce passes every chunk hop by hop, so its bus bandwidth is bounded by the slowest link in the ring — and almost every hop here crosses the network. Achievable is therefore about 50 GB/s per GPU, not 900.

Put a step on it. A 10-billion-parameter model in bfloat16 holds 20 GB of gradients. Ring all-reduce moves roughly twice the message size per participant, so about 40 GB per GPU per step.

  • Fragmented placement: 40 GB ÷ 50 GB/s ≈ 0.8 s of communication per step.
  • Eight whole nodes: the intra-node portion runs at NVLink speed and only the inter-node reduction crosses the wire, so the communication term falls by roughly an order of magnitude.

If compute is about 0.5 s per step, the fragmented job spends more than half its wall-clock waiting on the network while the scheduler reports it as fully allocated and the GPUs report high utilisation.

Which assumption dominates the error

Placement, by roughly an order of magnitude — not FLOPs. Second is the parallelism strategy: tensor and expert parallelism exchange activations every layer and must stay inside the NVLink domain, while pure data parallelism exchanges once per step and tolerates the network. A capacity model that asks only "how many GPUs" cannot distinguish these two jobs, and they differ by a factor of ten in what the same hardware delivers.

What the report should say

  • Free whole NVLink domains, not free GPUs. Ninety-six free GPUs at two per node is zero schedulable domains for a collective job.
  • Largest contiguous run of free domains, which is what bounds the biggest job you can admit.
  • A fragmentation ratio: free GPUs divided by free GPUs in whole domains. A ratio far above one is a scheduling defect, not a capacity shortage, and buying hardware will not fix it.

The decision rule follows: size and allocate accelerator capacity in units of the interconnect domain, with gang scheduling so a job gets whole domains or waits, and a defragmentation window where preemptible work is drained to re-form domains.

When this is the wrong answer

It flips entirely for jobs with no collectives. A fleet serving single-GPU inference replicas has no ring and no domain requirement, so free GPUs really is the right number and domain accounting is pure overhead. The same applies to embarrassingly parallel sweeps. Check the workload mix before imposing gang scheduling, because it costs real utilisation: whole domains held idle waiting for a job that fits.

Common weak answers

  • "Add more GPUs." The cluster is fully allocated. More hardware arrives in the same fragmented state unless the scheduler changes.
  • "Utilisation is high so capacity is fine." GPU utilisation counts cycles in kernels, and a kernel blocked inside a collective still counts. Utilisation cannot see this failure at all.