Skip to content

Chapter 12: Scale-Up Networking -- NVLink Switch System

HPC Networking FoundationsAdvanced45 min read

Every chapter so far has treated the DGX node as an atomic unit: eight GPUs wired together internally by NVLink, connected externally to the rest of the cluster by InfiniBand or RoCEv2. That external fabric -- the scale-out fabric -- is what Chapters 3 through 11 are about.

This chapter covers a different question: what happens when the scale-out fabric is not the bottleneck, but the boundary between scale-up and scale-out is?

The answer is the NVLink Switch System -- a way to extend NVLink beyond the chassis boundary and connect up to 32 DGX H100 nodes as if they were one enormous GPU.