Chapter 0 told you the three networks exist. It gave you the checkpoint math: a 70B parameter model produces a 140 GB checkpoint, written at 800 Gb/s, taking roughly 1.4 seconds. That number felt clean and manageable.
Here is what Chapter 0 did not tell you.
When 32 DGX nodes each attempt a checkpoint simultaneously -- which is exactly what happens at a synchronised save point in distributed training -- those 32 nodes produce 32 x 140 GB = 4.5 TB of write traffic in the same 1.4-second window. That traffic hits the storage fabric all at once. If the storage fabric cannot sustain 32 parallel writes at full speed, checkpoint time extends. Every second of extension is 32,000 GPU-hours sitting idle on a 4,000-node cluster at $500 per GPU-hour.
At $133 per second of cluster idle, a storage bottleneck that adds ten seconds to every checkpoint -- and checkpoints happen every ten to thirty minutes -- costs between $25,000 and $75,000 per day in wasted compute. Silently. With no error messages visible on the compute fabric.
The storage fabric is the network you cannot see failing from the perspective of the tools this course has taught you so far.
This chapter explains how it works, why it is engineered the way it is, and what to look for when it goes wrong.