Part II: Training and Inference Workloads

Training updates model parameters and keeps the activations, gradients and optimizer buffers needed for those updates. Inference uses the parameters to answer requests. It first processes each prompt, then generates further tokens while retaining attention keys and values from earlier tokens. These differences change both the memory required and the work that limits speed.

Chapter 7 covers training memory and numerical precision. Chapters 8-10 cover inference memory, token generation and request scheduling. Chapter 11 brings both workloads into repeatable performance experiments.

Training and inference form sibling branches after the foundations. Training and precision occupy one branch; inference workload, decode state, and serving occupy the other; both feed controlled experiments. Section labels identify part ii chapter map: 7 - Training and Precision; 8 - Inference Workloads; 9 - Decode and KV Cache; 10 - Serving Engines; 11 - Performance Experiments.
Figure 1: Training and serving keep different state live and therefore expose different limits.

Chapters in this part

  • 7. Training and Precision: Explain how training phases, live state, and numeric formats determine step time, memory pressure, and the correctness conditions for precision changes.
  • 8. Inference Workloads: Explain the inference software path, prefill and decode metrics, memory sizing, and compression choices that determine request capacity and latency.
  • 9. Decode and KV Cache: Explain why decode performance depends on cache growth, memory layout, batching, and kernel shape, then connect those mechanisms to measured throughput and latency.
  • 10. Serving Engines: Explain how a serving engine admits, batches, caches, and schedules requests, and how those choices trade throughput, latency, memory capacity, and fairness under a fixed workload.
  • 11. Performance Experiments: Turn performance questions into reproducible experiments with controlled inputs, correctness checks, warmup, profiling, and explicit interpretation.