Appendix A — Reference Material
Provide compact formulas, software-control ownership, practice checks, and source mappings for implementing and checking the methods taught in the chapters.
This appendix collects quantities and controls that are useful during implementation. The main chapters remain the source for mechanisms and trade-offs. The entries here provide a compact way to recover a formula, locate the software component that exposes a setting, or find supporting research and documentation.
A.1 Core formulas
These equations summarize the first-order models used throughout the guide. Their estimates depend on the measurement boundary: bytes refer to the memory tier being studied, numerical formats determine storage size, and theoretical peak values differ from achieved values.
The KV-cache product assumes equally long retained sequences, the same key/value-head dimensions in every counted layer, and separate cache storage for each sequence. Shared prefixes, sliding windows, and other cache representations require different counts. For more than one output token, TPOT is the mean interval after the first token. With one output token, end-to-end latency is approximately TTFT. Zero output tokens are outside this model. Ring traffic counts bytes sent by one rank, not sent and received bytes added together. The full-mesh expression counts one undirected link per pair of distinct endpoints. Activation storage is an approximation for the saved tensors of the particular training implementation.
\[ \text{AI} = \frac{\text{FLOPs}}{\text{bytes moved}} \tag{A.1}\]
\[ \text{achieved throughput} = \frac{\text{FLOPs}}{\text{seconds}} \tag{A.2}\]
\[ \text{roofline bound} = \min(\text{peak compute}, \text{AI} \times \text{peak bandwidth}) \tag{A.3}\]
\[ \text{AI}_{\text{ridge}} = \frac{\text{peak compute}}{\text{peak memory bandwidth}} \tag{A.4}\]
\[ \text{KV bytes} = 2 \times B \times T \times L \times H_{\text{kv}} \times d_h \times b \tag{A.5}\]
\[ \text{E2E latency} \approx \text{TTFT} + (\text{number of output tokens} - 1) \times \text{TPOT} \tag{A.6}\]
\[ \text{communication time} \approx \text{latency} + \frac{\text{bytes transferred}}{\text{effective bandwidth}} \tag{A.7}\]
\[ \text{ring all-reduce bytes sent per rank} \approx \frac{2(N - 1)}{N} \times \text{tensor}_{\text{bytes}} \tag{A.8}\]
\[ \text{full-mesh links} = \frac{N(N - 1)}{2} \tag{A.9}\]
\[ \text{pipeline bubble fraction} \approx \frac{K - 1}{M + K - 1} \tag{A.10}\]
\[ \text{training activation bytes} \approx B \times T \times d \times L \times k_{\text{act}} \times \text{bytes}_{\text{per element}} \tag{A.11}\]
\[ \text{expected speculative advance} = \sum_{i=0}^{\gamma} \alpha^i \tag{A.12}\]
\[ \text{speculative speedup estimate} = \frac{\sum_{i=0}^{\gamma} \alpha^i}{\gamma c + 1} \tag{A.13}\]
The symbols below use the same meanings and units as their first derivations in the main chapters.
| Symbol | Meaning and unit | Defined in |
|---|---|---|
| AI | Arithmetic intensity, measured in floating-point operations per byte moved at the stated memory tier | 5.3 |
| B | Batch size or number of active sequences, according to the formula | 7.2 and 8.6 |
| T | Sequence length in tokens (stored tokens per sequence in the KV-cache formula) | 7.2 and 8.6 |
| L | Number of model layers represented by the estimate | 7.2 and 8.6 |
| H_kv | Number of key-value heads per attention layer | 8.6 |
| d_h | Elements in one key or value head vector | 8.6 |
| b | Bytes per stored element | 7.2 and 8.6 |
| d | Model hidden dimension in elements | 7.2 |
| k_act | Implementation-dependent count of saved activation-sized tensors per layer | 7.2 |
| N | Number of communication ranks or topology endpoints | 12.6 |
| K | Number of pipeline stages | 13.5 |
| M | Number of micro-batches in the GPipe estimate | 13.5 |
| alpha | Average speculative-token acceptance probability | 10.6 |
| gamma | Number of draft tokens proposed per speculative iteration | 10.6 |
| c | Draft-step cost divided by target-step cost in the simplified model | 10.6 |
The pipeline expression is a balanced GPipe-style fill/drain estimate.1 The activation expression uses \(k_{\text{act}}\) as an architecture-dependent and implementation-dependent count of saved activation-sized tensors per layer. The speculative expressions assume constant independent acceptance and a simple sequential draft-cost model, with \(0 \leq \alpha \leq 1\) and \(c \geq 0\).2 Writing the geometric series as a finite sum also covers perfect acceptance: at \(\alpha=1\), expected advance is \(\gamma+1\) tokens and estimated speedup is \((\gamma+1)/(\gamma c+1)\). The corresponding chapters explain which assumptions matter for capacity and performance estimates.
Each formula answers a limited question. The roofline estimate compares arithmetic work with memory traffic. The KV-cache formula estimates storage. The latency and communication formulas estimate time for the stated workload. None accounts for every delay in a running system. Profiles reveal waiting between launches, competing work, and transfers that overlap with computation.
A.2 Software layers and controls
Each setting is interpreted by a particular component. Its effects may extend to other layers. The groups below identify the component that interprets each setting and the behavior it can change.
- PyTorch model and tensor layer: PyTorch defines tensors, modules, automatic differentiation, optimizers, and high-level device operations.
torch.cuda: Device, stream, event, memory-statistic, synchronization, and CUDA Graph APIs select devices and expose execution controls. Use them to establish which device and path ran.torch.profiler: Records framework operations and their associated CPU, CUDA, shape, and memory events.torch.compile(...): Requests compiled execution through the PyTorch API. TorchDynamo, AOTAutograd, and TorchInductor implement compatible parts of that request. Guards, graph breaks, and backend support determine the result.torch.distributed: Creates process groups and submits point-to-point or collective communication through a selected backend.
- Model and checkpoint layer: Model libraries construct tokenizers and architectures and interpret checkpoint metadata. Serialization libraries store tensor payloads.
- Transformers: Provides model configuration, tokenizer factories, checkpoint model classes, and generation baselines for supported architectures.
- safetensors: Stores tensor payloads with explicit metadata and avoids executable pickle payloads. Loading speed still depends on storage and placement.
- CUDA and kernel layer: The CUDA runtime and driver manage contexts, memory, streams, events, modules, and launches. Vendor libraries and generated kernels perform the device work.
- cuBLAS and cuBLASLt: Tuned dense linear-algebra kernels are selected according to dimensions, layout, dtype, alignment, and hardware support.
- cuDNN: Provides tuned implementations for supported neural-network operations.
- Triton: Expresses custom tile-oriented GPU kernels with program IDs, address calculations, masks, loads, computation, and stores.
CUDA_VISIBLE_DEVICES: Changes which GPUs a process can see. It does not move tensors or improve a kernel by itself.PYTORCH_CUDA_ALLOC_CONF: Changes selected PyTorch caching-allocator policies when allocation behavior or fragmentation is being investigated.
- Compiler and representation layer: Compiler components transform compatible tensor programs, while quantization tools prepare lower-precision representations that still require a supported execution path.
- TorchDynamo, AOTAutograd, and TorchInductor: These components capture Python/PyTorch regions, stage differentiable graphs, and lower them to generated kernels or library calls.
- bitsandbytes and torchao: Provide selected low-precision weight, optimizer, sparsity, or quantization paths. Exact model, format, GPU, and kernel support is version-dependent.
- Serving layer: A serving engine admits requests, schedules prefill and decode work, allocates KV-cache blocks, and records service metrics.
- vLLM and SGLang: Examples of serving engines with request scheduling, KV-cache management, and prefix-reuse capabilities whose exact interfaces change by release.
- TensorRT-LLM: A deployment stack for building and running supported optimized inference engines, including multi-GPU configurations.
max_num_batched_tokens: A vLLM-style token budget that limits how much work the scheduler admits in an iteration.3--gpu-memory-utilization: A vLLM-style reservation target used when dividing device capacity among model state, workspaces, and KV cache.--max-model-len: A vLLM-style context limit that changes the worst-case KV-cache requirement.--kv-cache-dtype: Selects KV-cache storage precision only when the engine and attention kernels support that format.
- Communication layer: NVIDIA Collective Communications Library (NCCL) selects collective algorithms and transports for CUDA tensors on NVIDIA systems. The physical path still depends on topology and configuration.
NCCL_DEBUG=INFO: Records initialization, topology, algorithm, and transport information useful for diagnosis.NCCL_SOCKET_IFNAME: Selects the network interface used by NCCL socket transports.NCCL_IB_DISABLE: Disables InfiniBand/RDMA paths for a controlled fallback comparison.NCCL_P2P_DISABLE: Disables direct peer-to-peer GPU paths for a controlled topology comparison.
- Measurement layer: Profiler selection follows the object under inspection: PyTorch operations, gaps in the CPU/GPU timeline, or the instructions and memory accesses of one kernel. The tools below provide these different views. Perfetto is a viewer for exported traces.
- PyTorch Profiler: Maps framework operations to CPU and CUDA activity and can export a trace.
- Nsight Systems: Shows process, CUDA, copy, kernel, and collective timing on a system timeline.
- Nsight Compute: Inspects one GPU kernel’s instructions, occupancy, memory transactions, and stalls.
- Perfetto: Displays exported Chrome-trace timelines for interactive inspection.
Version, GPU architecture, dtype, shape, and topology can change whether a control is supported or useful. Logs and traces establish which kernel or transport actually ran. A setting alone does not establish its effect. The corresponding explanations and source notes are in Chapter 3, Chapter 6, Chapter 10, and Chapter 12.
A.3 Compatibility checks
A package can install successfully and the program can run without using the expected optimized kernel. The attention kernel, compiled region, or communication path that actually ran is established from runtime evidence. The runtime may have selected a slower supported alternative.
Checklist: Compatibility evidence
- Hardware and driver: The record contains the GPU model, memory capacity, driver, framework-visible CUDA runtime, and interconnect topology.
- Framework and kernel stack: The record contains PyTorch, Triton, attention backend, Transformers, quantization library, and serving-engine versions.
- Attention path: Support covers dtype, head dimension, causal or other masks, grouped-query layout, sequence length, and training or inference mode.
- Quantization path: The weight and KV formats have kernels for the target GPU, and accuracy is evaluated on the intended workload.
- Compiler and graph path: Keep compilation and warmup out of steady-state timing. Establish graph breaks, recompilations, and fixed-address requirements before CUDA Graph capture.
- Communication path: NCCL or another backend selects the intended local or network transport, and profiler traces agree with the assumed topology.
A compatibility check is complete only when the observed execution path matches the intended path. Package versions are necessary evidence, but kernel names, compiler logs, memory accounting, and communication traces show what actually ran.
A.4 Practice checks
The twelve questions below mix numerical estimates, conceptual explanations, and code reading. Work each problem without the answer key, state assumptions and units, and name the measurement that would confirm the result.
- P1 KV-cache budget: For 4 active sequences, 2,048 stored tokens, 32 layers, 8 key-value heads, head dimension 128, and 2-byte elements, compute the raw KV-cache size in GiB.
- P2 Model-scale units: Using the Llama 3.3 70B configuration in Section 8.6, compare one sequence at 32,000 tokens with one sequence at 32,768 tokens. Report both cache sizes in GiB. Explain why writing only “32K tokens” can make the estimate ambiguous, and distinguish GB from GiB.
- P3 Roofline: A kernel performs 2e12 FLOPs, moves 1e12 bytes, and runs in 0.020 seconds on a device rated at 200 TFLOP/s and 3 TB/s. Compute arithmetic intensity, the memory ceiling, and achieved throughput. Then explain the result.
- P4 Ring all-reduce: Eight ranks run ring all-reduce on a 1 GiB tensor. Use the first-order traffic formula to estimate bytes sent per rank, then list two reasons a trace can differ from the estimate.
- P5 Pipeline bubble: A pipeline has four stages and eight micro-batches. Use the bubble-fraction formula to estimate the idle fraction before considering load imbalance or communication.
- P6 Prefill and decode: Why can transformer projections be GEMM-like during prefill and GEMV-like during small-batch decode? Which metric may improve and which may worsen when batching increases?
- P7 FlashAttention: Which mathematical result does FlashAttention preserve, which intermediate does it avoid materializing in HBM, and which performance resource does it target? Which serving concern remains unresolved?
- P8 PagedAttention: A request maps logical KV blocks [0, 1, 2] to physical blocks [7, 2, 15]. Explain what remains ordered, what becomes non-contiguous, and why this helps admission capacity.
- P9 Software ownership: For each setting, identify the component that interprets it and describe the work or data it can change:
torch.compile(...),--kv-cache-dtype,tl.load(..., cache_modifier=...), andNCCL_DEBUG=INFO. - P10 Speculative decoding: Give two conditions that can reduce latency and two conditions that can erase the gain. Include target-model verification and rejected-token correction.
- P11 Triton GEMV: In the GEMV template in Section 11.2.4, explain what
tl.program_id(0), the maskedtl.load, the reduction, andtl.storeeach control. Identify one layout or tile choice that could make the kernel slower. - P12 Distributed Data Parallel: In the DDP template, explain the roles of
DistributedSampler,loss.backward(), and NCCL. Identify one single-GPU cause and one cross-rank cause of a slow step.
The answer key is not a replacement for measurement. Several answers are first-order diagnoses, so a real experiment still tests cache effects, overlap, fragmentation, graph breaks, or topology.
A.5 Answer key
The answers follow P1 through P12 in order. A different answer can still be correct when it states its assumptions, units, and supporting measurement clearly.
- P1 KV-cache budget: The raw size is 1,073,741,824 bytes, exactly 1 GiB. This counts separate, equally long caches for all four sequences. Block metadata, fragmentation, and workspaces require additional memory. Prefix sharing is excluded.
- P2 Model-scale units: At 32,000 tokens the result is 10,485,760,000 bytes, about 9.77 GiB. At 32,768 tokens it is 10,737,418,240 bytes, exactly 10 GiB. State whether K means decimal thousands or 1,024-token units.
- P3 Roofline: Arithmetic intensity is 2 FLOPs per byte. The memory ceiling is 6 TFLOP/s, but the stated runtime implies 100 TFLOP/s. The values contradict the simple model, so recheck the byte count, cache assumption, FLOP count, and timing boundary.
- P4 Ring all-reduce: Each rank sends 1.75 GiB and receives 1.75 GiB. The formula counts one direction, not the sum of sent and received bytes. Algorithm choice, chunking, topology, protocol overhead, overlap, contention, and rank skew can change a trace.
- P5 Pipeline bubble: The first-order idle fraction is 3/11, or about 27.3 percent. Stage imbalance, communication, and scheduling overhead change the measured result.
- P6 Prefill and decode: Prefill exposes many token positions and reuses weights across larger matrix operations, so it can raise compute use and mainly affects TTFT. Small-batch decode adds one token per request while rereading weights and KV history. Batching can improve throughput while increasing queueing or TPOT.
- P7 FlashAttention: It preserves the attention result while avoiding a full quadratic score or probability matrix in HBM. It reduces off-chip memory traffic. It does not allocate KV blocks, admit requests, or choose a serving schedule.
- P8 PagedAttention: Logical token order remains [0, 1, 2], while physical blocks can be [7, 2, 15]. On-demand fixed-size allocation reduces reserved waste and fragmentation at the cost of block-table lookup and scheduler/kernel complexity.
- P9 Software ownership: The PyTorch compiler/runtime owns
torch.compile(...). The serving engine owns--kv-cache-dtype. Triton kernel code ownstl.load(..., cache_modifier=...), and NCCL ownsNCCL_DEBUG=INFO. A setting is interpreted by a specific component, but its effects may extend to other parts of the system. For example, a serving engine’s cache precision setting changes how many bytes the cache occupies and may change which attention kernel can run. Use logs or traces to confirm the actual effect rather than assuming it from the setting’s name. - P10 Speculative decoding: It helps when the draft is cheap, acceptance is high, and target verification is efficient. Low acceptance, a costly draft, a saturated target batch, or rollback, cache, sampling, and scheduling overhead can erase the gain.
- P11 Triton GEMV: The program ID selects an output row, masked loads protect the final partial tile, the reduction accumulates the row dot product, and the store writes one output. Poor coalescing, tiles that are too large, register spilling, or too little parallelism can slow it down.
- P12 Distributed Data Parallel: The sampler partitions input across ranks, backward makes gradients ready, and NCCL performs the collective reduction. A slow local kernel is a single-GPU cause. Topology contention, rank skew, or exposed all-reduce time are cross-rank causes.
A.6 Source map
The source map records where lecture, summary, notebook, and externally researched material contributes to the guide. A source unit is a notebook cell, a PDF page, a paragraph in a summary, or a whole document where the source has not been divided further. Notebook exercises are integrated into the chapters that teach roofline reasoning, decode and cache behavior, compilation, CUDA Graphs, and experimental practice rather than being isolated as homework appendices.
| Source | Mapped units | Coverage status | Guide sections |
|---|---|---|---|
| [Py4DP-L3] [Optional] PyTorch on GPU.ipynb | cell 1-58 (58 mapped) | 58 cells combined into the explanation | 3.1 |
| [Py4DP-L3] Gradients & Logistic regression in PyTorch.ipynb | cell 1-56 (56 mapped) | 56 cells combined into the explanation | 1.4 |
| [Py4DP-L3] PyTorch basics.ipynb | cell 1-78 (78 mapped) | 78 cells combined into the explanation | 1.1 |
| 04 _ Performance Engineering - Week 01 _ Systems, GPUs & Networking - Summary.docx | unit 1-82 (74 mapped) | 70 units covered; 4 omitted | 1.6, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 3.2, 3.3, 3.5, 3.6, 5.1, 5.2, 5.3, 5.4, 5.4.3, 5.5, 8.2, 8.3, 11.1, 11.1.6, 12.3, 12.5, 12.12 |
| 04 _ Performance Engineering - Week 02 _ Efficient Inference- Architecture, Quantization & Compilers - Summary.docx | unit 1-57 (53 mapped) | 51 units covered; 2 omitted | 1.6, 8.2, 8.3, 8.6, 9.4, 9.5, 10.2, 10.2.1, 10.2.2 |
| hw1_roofline.ipynb | unit 1-18 (18 mapped) | 18 units covered | 5.1, 5.2, 5.3, 5.4, 11.1, 11.2 |
| hw2_decode_optimization.ipynb | unit 1-17 (17 mapped) | 17 units covered | 9.1, 9.2, 9.3, 11.1, 11.2 |
| hw3_compile_cuda_graphs.ipynb | unit 1-21 (21 mapped) | 21 units covered | 6.1, 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 11.1, 11.2 |
| lecture 1 - foundations of supervised training [Autosaved].pptx.pdf | unit 1-35 (35 mapped) | 14 units covered; 21 omitted | 1.4, 1.5, 8.5, 13.4 |
| lecture 2 - Data Parallel Training (DDP) [Autosaved].pptx.pdf | unit 1-41 (41 mapped) | 32 units covered; 9 omitted | 11.1, 11.1.9, 12.6, 12.8, 12.8.1, 12.9, 13.1, 13.2, 13.3 |
| lecture 3 - Modern Large-Model Training Techniques.pptx.pdf | unit 1-61 (61 mapped) | 34 units covered; 27 omitted | 2.7, 8.4, 8.5, 12.8, 13.4, 13.5, 13.5.1, 13.5.2 |
| Performance Engineering - L1.pdf | unit 1-89 (89 mapped) | 83 units covered; 6 omitted | 1.6, 2.1, 2.2, 2.3, 2.4, 2.5, 2.6, 3.2, 3.3, 3.5, 3.6, 5.1, 5.2, 5.3, 5.4, 5.4.3, 5.5, 8.2, 8.3, 11.1, 11.1.6, 12.3, 12.5, 12.12 |
| Performance Engineering - L2-2.pptx.pdf | unit 1-139 (139 mapped) | 133 units covered; 6 omitted | 6.2, 6.3, 6.4, 6.5, 6.6, 6.7, 6.7.2, 8.2, 8.3, 8.6, 9.4, 9.5, 9.6, 10.1, 10.2, 10.2.1, 10.2.2, 10.3, 10.4, 10.5, 10.5.1, 10.5.2, 10.5.3, 10.7, 10.8, 10.9, 12.2, 12.5, 13.7, A.2 |
| Performance Engineering - L4.pdf | unit 1-61 (61 mapped) | 56 units covered; 5 omitted | 3.2, 3.5, 3.6, 3.7, 4.1, 4.1.1, 4.2, 4.3, 4.5, 4.6, 4.7, 5.4, 6.4, 11.2, 11.2.3, 11.2.5 |
| Performance Engineering - L6.pdf | unit 1-81 (81 mapped) | 79 units covered; 2 omitted | 7.4, 8.7.1, 8.7.2, 8.7.4, 9.6, 10.4, 10.5, 10.5.1, 10.5.2, 10.5.3, 10.7, 10.8, 10.9, 12.2, 12.5, 13.7 |
| solution_1.1.ipynb | unit 1-13 (13 mapped) | 11 units covered; 2 omitted | 10.6, 10.6.1, 11.2, 11.2.10 |
| solution_1.4.ipynb | unit 1-28 (28 mapped) | 27 units covered; 1 omitted | 10.6, 10.6.1, 11.2, 11.2.10, 11.2.11 |
| Speculative Decoding Paper Summary.docx | unit 1-58 (54 mapped) | 45 units covered; 9 omitted | 8.2, 9.1, 10.6, 10.6.1, 11.2, 11.2.10 |
| Speculative Decoding Workshop - Nebius Academy 260619.pptx.pdf | unit 1-25 (25 mapped) | 9 units covered; 16 omitted | 10.6, 10.6.1, 11.2, 11.2.10 |
| workshop.ipynb | unit 1-8 (8 mapped) | 7 units covered; 1 omitted | 10.6, 10.6.1, 11.2, 11.2.10 |
| master_outline.txt | document 1 | 1 omitted | No mapped guide section; see source notes |
The source map keeps one record for every slide, page, notebook cell, or source paragraph, including where it is used and why it was included, combined, treated as a duplicate, or omitted. The table groups those records by source so the course material remains easy to trace without repeating the full internal record in the book. The selected external sources below connect major methods to their teaching sections. The References section contains the complete list of cited works.
| Primary reference | Date | Guide section | URL |
|---|---|---|---|
| Dao et al. (2022) FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness | 2022-05-27 | 4.8 | https://arxiv.org/abs/2205.14135 |
| Kwon et al. (2023) Efficient Memory Management for Large Language Model Serving with PagedAttention | 2023-09-12 | 9.5 | https://arxiv.org/abs/2309.06180 |
| Zheng et al. (2024) SGLang: Efficient Execution of Structured Language Model Programs | 2024-06-06 | 10.5 | https://arxiv.org/abs/2312.07104 |
| Leviathan et al. (2023) Fast Inference from Transformers via Speculative Decoding | 2023-05-18 | 10.6 | https://arxiv.org/abs/2211.17192 |
| Rajbhandari et al. (2020) ZeRO: Memory Optimizations Toward Training Trillion Parameter Models | 2020-05-13 | 13.4 | https://arxiv.org/abs/1910.02054 |
| Shoeybi et al. (2019) Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism | 2019-09-17 | 13.5 | https://arxiv.org/abs/1909.08053 |
| Huang et al. (2019) GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism | 2019-07-25 | 13.5 | https://arxiv.org/html/1811.06965v5 |
| Williams et al. (2009) Roofline: An Insightful Visual Performance Model for Multicore Architectures | 2009-04-01 | 5.3 | https://dl.acm.org/doi/10.1145/1498765.1498785 |
| Tillet et al. (2019) Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations | 2019-06-22 | 4.2 | https://dl.acm.org/doi/10.1145/3315508.3329973 |
Software interfaces and hardware behavior can change after the dates shown here. Recheck the installed version and official documentation before treating a setting, backend, or limit as current operational fact.
A.7 Techniques by lever class
Section 1.6.1 names seven optimization levers, meaning seven kinds of change a technique can make. The table groups the guide’s techniques by the lever each one pulls, so that a measured bottleneck can be turned into a short list of candidate changes. The entries highlight the main levers discussed in the linked sections. A technique can affect more than one.
| Lever | Technique | Section |
|---|---|---|
| Do less work | Key-value cache | 9.1 |
| Do less work | Prefix caching and radix-style reuse | 10.5 |
| Do less work | Grouped-query attention | 8.6 |
| Do less work | Pruning and structured sparsity | 8.7.5 |
| Do less work | Distillation | 8.7.6 |
| Do less work | Sliding-window attention and attention sinks | 9.7 |
| Move fewer bytes | Kernel fusion | 4.6 |
| Move fewer bytes | Tiled attention, such as FlashAttention | 4.8 |
| Move fewer bytes | Memory coalescing | 3.6 |
| Move fewer bytes | Tile order and cache hints | 4.4.2 |
| Move fewer bytes | Weight quantization | 8.7.1 |
| Move fewer bytes | Key-value cache quantization and latent cache | 9.7 |
| Use faster hardware paths | Tensor Cores and shape alignment | 2.4 |
| Use faster hardware paths | Vendor kernel libraries | 3.1.2 |
| Use faster hardware paths | Mixed-precision training | 7.4 |
| Use faster hardware paths | Compiled inference engines | 10.9.2 |
| Keep the hardware busy | Continuous batching | 10.2 |
| Keep the hardware busy | Speculative decoding (parallel target verification) | 10.6 |
| Keep the hardware busy | Chunked prefill | 10.3 |
| Keep the hardware busy | Streams and copy-compute overlap | 3.3 |
| Keep the hardware busy | Communication and computation overlap | 12.9 |
| Keep the hardware busy | Gradient bucket overlap in DDP | 13.3 |
| Cut fixed overhead | torch.compile capture and fusion | 6.2 |
| Cut fixed overhead | CUDA Graph capture and replay | 6.5 |
| Cut fixed overhead | Scheduler work off the decode path | 10.2.2 |
| Fit in memory | Activation checkpointing | 7.2 |
| Fit in memory | PagedAttention and block tables | 9.5 |
| Fit in memory | ZeRO and FSDP sharding | 13.4 |
| Fit in memory | Offload to CPU memory | 13.4.1 |
| Fit in memory | External key-value cache layers | 10.5.3 |
| Scale out | Data parallelism | 13.2 |
| Scale out | Tensor parallelism | 13.5.1 |
| Scale out | Pipeline parallelism | 13.5 |
| Scale out | Sequence and context parallelism | 13.5.4 |
| Scale out | Expert parallelism and mixture-of-experts serving | 13.7.2 |
| Scale out | Separate prefill and decode workers | 10.4 |
A lever names what a change does, not how much it wins. What any of them can achieve for the whole run is bounded by the share of time the changed region holds, by Amdahl’s law in Section 1.6.1.
Huang, Y., et al. (2019). GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism. arXiv:1811.06965v5, Section 2.3 and Figure 2. https://arxiv.org/html/1811.06965v5. The bubble estimate assumes balanced stages. Communication delays and stage imbalance can increase measured idle time.↩︎
Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast inference from Transformers via speculative decoding. Proceedings of the 40th International Conference on Machine Learning, PMLR 202, 19274-19286. https://proceedings.mlr.press/v202/leviathan23a.html. The analysis assumes independent acceptance and a simplified relative draft cost. The formulas do not by themselves predict latency under a loaded serving scheduler.↩︎
vLLM contributors. (n.d.). Engine Arguments (v0.6.4.post1 documentation). https://docs.vllm.ai/en/v0.6.4.post1/models/engine_args.html. Documents the listed token, memory, context-length, and cache-format controls. Defaults and supported combinations can change by release.↩︎