Congestion control in AI fabrics is an algorithm-level competition measured in RTTs. The key insight across all five algorithms:
- DCQCN: 2-RTT feedback via ECN → CNP → rate reduce. Deployed on every ConnectX-7. Requires switch ECN thresholds and CNP DSCP trust to be correctly configured.
- Swift: 1-RTT feedback via RTT measurement at NIC. Excellent under incast. Requires custom NIC hardware. Not available on ConnectX-7.
- HPCC: 1-RTT feedback via INT headers from every switch hop. Most precise. Requires INT-capable switch ASIC. High operational complexity.
- TIMELY: Reacts to RTT gradient, not absolute RTT. Works on commodity NICs. Practical only when baseline RTT >> 20 µs (noise floor constraint).
- UEC CC: 1-RTT feedback via C-flag in ACK. No separate CNP. Optional NPM telemetry. Available today on AMD Pensando Salina; NVIDIA future roadmap.
For the overwhelming majority of deployed NVIDIA GPU clusters today: get DCQCN right. The difference between a correctly tuned DCQCN deployment and a misconfigured one is 20–40% JCT improvement. Labs 12 and 13 give you hands-on practice diagnosing the most common DCQCN misconfiguration patterns and benchmarking algorithm behaviour directly.