Introduction and Learning Path

Establish the audience, reading path, and measurement-first method that connect single-GPU execution, training and inference, and distributed systems.

Large language model performance is determined by more than the number of floating-point operations. Each training step or inference request creates tensors with different sizes, lifetimes, reuse patterns, and placement constraints. The hardware must move those tensors through a memory hierarchy, execute kernels, and sometimes exchange partial results with other GPUs. A useful optimization starts by identifying the active work and the resource that limits it.

Basic Python and PyTorch familiarity is enough to begin. Chapter 1 introduces the tensors used in training and generation. Chapters 2-4 explain where those tensors are stored and how GPU programs process them. Chapters 5-6 show how to measure that work and reduce compiler and runtime overhead. Equations and code examples support the memory estimates, timing comparisons, and implementations along the way.

Part I covers execution on one GPU. Part II applies those foundations to training and inference serving. Part III examines how GPUs share work and exchange data.

Thirteen chapters arranged as a reading path across three parts: single-GPU foundations, training and inference workloads, and distributed systems, with measurement loops returning to earlier decisions.
Figure 1: Book structure and chapter path

The chapter path is a reading order, not one runtime pipeline. Measurement recurs throughout the guide because a change should be checked in the phase and system where it is expected to help before it is trusted.

After the common foundations, local acceleration, inference serving, and distributed systems can be studied as separate tracks. Each track adds different limits and control points.

Foundations split into three parallel study tracks for local acceleration, inference serving, and distributed systems; the tracks merge into hands-on diagnosis and a measure-again feedback loop.
Figure 2: Study route from foundations to diagnosis

The tracks meet again in hands-on diagnosis. A reader can follow one track for immediate work, but reliable optimization still uses the same cycle: measure, change one relevant control, verify the result, and measure again.

What mathematical objects become in a running system

A tensor is an array of numbers, but running an operation also requires storing those numbers and moving them to the processor. Two implementations can compute the same result while moving different amounts of data. For example, reusing weights already available to a group of operations can reduce repeated memory reads. The guide connects the mathematics of each operation to these storage and execution costs.

Questions for each performance problem

For a slow training step or generation request, work through these questions:

  • Which part takes the time: input preparation, GPU computation, token generation, or communication?
  • What data does that part read, write, or keep in memory?
  • Is it doing arithmetic, moving data, or waiting? Use measurements to distinguish these possibilities.
  • Which code, library, or setting can change that work?
  • After the change, are the results still correct, is model quality acceptable, and has the original time or throughput target improved under the same workload?

What each main mechanism changes

Some optimizations reduce the work done for each token. Others reduce data transfers or keep the GPU from waiting. The right choice depends on which cost dominates the workload being measured.