Advanced LLM Performance Engineering

Contents
Front matter
- Introduction and Learning Path
Establish the audience, reading path, and measurement-first method that connect single-GPU execution, training and inference, and distributed systems.
Part I: Fundamentals
Connect model workloads to GPU hardware, CUDA execution, Triton kernels, measurement, compilation, and single-worker runtime behavior.
- Chapter 1: LLM Workloads and Performance Limits
Connect workload phases, live tensor state, resource limits, measurements, and software controls so performance work begins with a measurable bottleneck and its responsible layer. - Chapter 2: GPU Hardware and Memory
Explain how GPU execution units, the memory hierarchy, and interconnects determine the cost of moving and processing model tensors. - Chapter 3: Software to CUDA
Connect framework calls and settings to CUDA launches, streams, memory transfers, and the software layer that owns each control. - Chapter 4: Triton and Custom Kernels
Explain how Triton maps block-level programs to GPU work and when a measured layout or fusion problem justifies a custom kernel. - Chapter 5: Performance Models and Profiling
Explain how roofline models and profiler traces identify the limiting resource in a measured region, then guide a focused performance experiment. - Chapter 6: Compilation and Runtime
Compare compilation and CUDA Graph replay by the host or device work they change, their prerequisites, and the evidence needed to select either path.
Part II: Training and Inference Workloads
Apply the single-worker foundations to training, inference, decode, cache management, serving engines, and controlled performance experiments.
- Chapter 7: Training and Precision
Explain how training phases, live state, and numeric formats determine step time, memory pressure, and the correctness conditions for precision changes. - Chapter 8: Inference Workloads
Explain the inference software path, prefill and decode metrics, memory sizing, and compression choices that determine request capacity and latency. - Chapter 9: Decode and KV Cache
Explain why decode performance depends on cache growth, memory layout, batching, and kernel shape, then connect those mechanisms to measured throughput and latency. - Chapter 10: Serving Engines
Explain how a serving engine admits, batches, caches, and schedules requests, and how those choices trade throughput, latency, memory capacity, and fairness under a fixed workload. - Chapter 11: Performance Experiments
Turn performance questions into reproducible experiments with controlled inputs, correctness checks, warmup, profiling, and explicit interpretation.
Part III: Communication and Distributed Systems
Extend the same performance reasoning to GPU communication, collective operations, distributed training, and distributed serving.
- Chapter 12: GPU Communication
Explain why GPU communication enters the critical path, how topology and collective algorithms shape its cost, and how to measure it. - Chapter 13: Distributed Training and Serving
Compare distributed training and serving strategies by the object each divides, the communication each adds, and the memory, topology, and user-metric limits each must satisfy.
Appendices
- Appendix A. Reference Material
Provide compact formulas, software-control ownership, practice checks, and source mappings for implementing and checking the methods taught in the chapters. - References
Consolidate the primary and supporting sources cited for technical claims, interfaces, measurements, and stated limits throughout the guide.