Advanced LLM Performance Engineering

Author

Vitaly Rubinovich

Compilation date
Publisher ID
study-guide-publisher@0.8.14+873a57a

A cutaway GPU on an engineering workbench shows a compute die, memory modules, incoming token tiles, layered persistent state, a fast interconnect to a second GPU, and a timing instrument.

Contents

Front matter

  • Introduction and Learning Path
    Establish the audience, reading path, and measurement-first method that connect single-GPU execution, training and inference, and distributed systems.

Part I: Fundamentals

Connect model workloads to GPU hardware, CUDA execution, Triton kernels, measurement, compilation, and single-worker runtime behavior.

Part II: Training and Inference Workloads

Apply the single-worker foundations to training, inference, decode, cache management, serving engines, and controlled performance experiments.

  • Chapter 7: Training and Precision
    Explain how training phases, live state, and numeric formats determine step time, memory pressure, and the correctness conditions for precision changes.
  • Chapter 8: Inference Workloads
    Explain the inference software path, prefill and decode metrics, memory sizing, and compression choices that determine request capacity and latency.
  • Chapter 9: Decode and KV Cache
    Explain why decode performance depends on cache growth, memory layout, batching, and kernel shape, then connect those mechanisms to measured throughput and latency.
  • Chapter 10: Serving Engines
    Explain how a serving engine admits, batches, caches, and schedules requests, and how those choices trade throughput, latency, memory capacity, and fairness under a fixed workload.
  • Chapter 11: Performance Experiments
    Turn performance questions into reproducible experiments with controlled inputs, correctness checks, warmup, profiling, and explicit interpretation.

Part III: Communication and Distributed Systems

Extend the same performance reasoning to GPU communication, collective operations, distributed training, and distributed serving.

  • Chapter 12: GPU Communication
    Explain why GPU communication enters the critical path, how topology and collective algorithms shape its cost, and how to measure it.
  • Chapter 13: Distributed Training and Serving
    Compare distributed training and serving strategies by the object each divides, the communication each adds, and the memory, topology, and user-metric limits each must satisfy.

Appendices

  • Appendix A. Reference Material
    Provide compact formulas, software-control ownership, practice checks, and source mappings for implementing and checking the methods taught in the chapters.
  • References
    Consolidate the primary and supporting sources cited for technical claims, interfaces, measurements, and stated limits throughout the guide.