Part III: Communication and Distributed Systems

Multiple GPUs can add capacity or throughput when one device is insufficient. Whether they improve latency depends on how the work is divided and on the communication it requires. The required exchanges of model state or partial results can place communication on the critical path.

When a model is spread across GPUs, a step may have to wait for data from another GPU. The delay depends on how much data is sent, which links carry it, and whether useful computation can continue during the transfer. Chapter 12 explains those transfers and how to measure them. Chapter 13 compares ways to divide model data and computation across GPUs.

Two ranks exchange data. A timing sketch shows the early rank waiting before an aligned exchange. Communication costs inform layouts illustrated by identical replicas, complementary state shards and successive layer groups. Serving compute reads and appends KV state and produces a token. A feedback arrow calls for measuring the chosen layout.
Figure 1: Communication costs constrain distributed layouts. The timing sketch shows one qualitative readiness wait before a common exchange; real collectives can progress differently. Replicas, disjoint state shards, and pipeline-style layer partitions illustrate different choices, while serving retains KV state beside computation.

Chapters in this part

  • 12. GPU Communication: Explain why GPU communication enters the critical path, how topology and collective algorithms shape its cost, and how to measure it.
  • 13. Distributed Training and Serving: Compare distributed training and serving strategies by the object each divides, the communication each adds, and the memory, topology, and user-metric limits each must satisfy.