Part I: Fundamentals

LLM workloads can appear slow for several different reasons: a tensor may not fit in memory, bytes may arrive too slowly, kernels may not expose enough parallel work, or the host may leave the GPU idle. The first task is to identify the active workload phase and the physical resource that limits it.

Chapter 1 describes the data needed during training and generation. Chapters 2-4 show where it is stored and how CUDA and Triton kernels process it. Chapter 5 uses measurements to distinguish arithmetic time, memory traffic and waiting. Chapter 6 examines overhead from dispatching and launching repeated work.

Single-GPU foundations connect live workload state to nested GPU hardware, CUDA and Triton execution, roofline and profiler measurements, compilation, and CUDA Graph replay. Section labels identify part i chapter map: 1 - LLM Workloads and Performance Limits; 2 - GPU Hardware and Memory; 3 - Software to CUDA; 4 - Triton and Custom Kernels; 5 - Performance Models and Profiling; 6 - Compilation and Runtime.
Figure 1: Workload phases meet GPU hardware, programming models, and measurable performance limits.

The active phase determines which objects are live and which operations run. Those operations pass through software and hardware before measurements can distinguish compute, memory, and launch limits. Data movement and reuse depend on the selected operation and kernel rather than one mandatory path.

Chapters in this part

  • 1. LLM Workloads and Performance Limits: Connect workload phases, live tensor state, resource limits, measurements, and software controls so performance work begins with a measurable bottleneck and its responsible layer.
  • 2. GPU Hardware and Memory: Explain how GPU execution units, the memory hierarchy, and interconnects determine the cost of moving and processing model tensors.
  • 3. Software to CUDA: Connect framework calls and settings to CUDA launches, streams, memory transfers, and the software layer that owns each control.
  • 4. Triton and Custom Kernels: Explain how Triton maps block-level programs to GPU work and when a measured layout or fusion problem justifies a custom kernel.
  • 5. Performance Models and Profiling: Explain how roofline models and profiler traces identify the limiting resource in a measured region, then guide a focused performance experiment.
  • 6. Compilation and Runtime: Compare compilation and CUDA Graph replay by the host or device work they change, their prerequisites, and the evidence needed to select either path.