Skip to content

Spectrum-X Architecture and the AI Factory Platform · Part 9 of 9

Chapter Summary

Spectrum-X is not merely a new switch model — it is a platform-level design that must be understood as the integration of four components working together. The Spectrum-4 ASIC provides the hardware adaptive routing engine and deep 48 MB packet buffer that absorb incast bursts and eliminate hash collision-driven load imbalance. The BlueField-3 SuperNIC (DGX B200 only) provides the in-NIC reorder buffer that decouples per-packet adaptive routing from the RoCEv2 Go-Back-N retransmission protocol — without BF3, aggressive adaptive routing degrades throughput rather than improving it. The DOCA SDK provides programmable packet processing from the host Linux environment. NetQ provides the fabric-wide push-based telemetry that makes Spectrum-X operationally manageable at scale.

The Scalable Unit is the deployment atom. Understand the rail-wired topology, the failure domain properties it produces, and the BasePOD/SuperPOD hierarchy before attempting any production deployment. Storage tier design is inseparable from network design: GPUDirect Storage, dedicated storage leaf switches, and NVMe-oF over RDMA are first-class networking concerns, not afterthoughts.

Chapter 25 builds directly on this foundation, covering RoCEv2 PFC/ECN configuration, DCQCN parameter tuning, and the specific NVUE command sequences required to bring a Scalable Unit's training fabric into production-ready state. Lab 16 provides hands-on practice with the physical bring-up audit procedure that must precede any RoCE configuration.

Key numbers to keep handy for platform planning:

  • Spectrum-4 aggregate bandwidth: 51.2 Tbps
  • Spectrum-4 shared buffer: 48 MB (vs Tomahawk 4: 12 MB)
  • Cut-through latency: <300 ns
  • BF3 ARM cores: 16× Cortex-A78AE
  • Standard SU composition: 8 DGX + 2 SN5600 + 1 SN4600C
  • BasePOD = 4 SUs = 32 DGX nodes
  • BFD detection timer: 50 ms × 3 = 150 ms MTTD