References

Consolidate the primary and supporting sources cited for technical claims, interfaces, measurements, and stated limits throughout the guide.

Sources are numbered by first citation in the book. Footnotes identify the supported claims and their limits. Software documentation describes the cited interface or version. It does not establish a general performance advantage.

  1. Grattafiori, A., et al. (2024). The Llama 3 Herd of Models. arXiv:2407.21783 (v1 2024-07-31, v3 2024-11-23).
  2. Touvron, H., et al. (2023). LLaMA: Open and Efficient Foundation Language Models. arXiv:2302.13971.
  3. Holtzman, A., Buys, J., Du, L., Forbes, M., & Choi, Y. (2020). The Curious Case of Neural Text Degeneration. ICLR 2020. arXiv:1904.09751.
  4. PyTorch. (2024). Tensor Attributes (PyTorch 2.4 documentation). https://docs.pytorch.org/docs/2.4/tensor_attributes.html.
  5. Kingma, D. P., & Ba, J. (2015). Adam: A Method for Stochastic Optimization. 3rd International Conference on Learning Representations. https://arxiv.org/abs/1412.6980.
  6. Rajbhandari, S., Rasley, J., Ruwase, O., & He, Y. (2020). ZeRO: Memory Optimizations Toward Training Trillion Parameter Models. https://arxiv.org/abs/1910.02054 (v3 2020-05-13).
  7. Ainslie, J., et al. (2023). GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. EMNLP 2023. arXiv:2305.13245.
  8. Kaplan, J., et al. (2020). Scaling Laws for Neural Language Models. arXiv:2001.08361.
  9. PyTorch. (2024). CrossEntropyLoss (PyTorch 2.4 documentation). Input and target shapes.
  10. NVIDIA. (2024). CUDA C++ Programming Guide 12.4. https://docs.nvidia.com/cuda/archive/12.4.0/cuda-c-programming-guide/index.html.
  11. Amdahl, G. M. (1967). Validity of the single processor approach to achieving large scale computing capabilities. In Proceedings of the April 18-20, 1967, Spring Joint Computer Conference (AFIPS ’67 Spring), 483-485. https://doi.org/10.1145/1465482.1465560.
  12. NVIDIA. (n.d.). Parallel Thread Execution ISA 8.4 (CUDA 12.4 documentation). https://docs.nvidia.com/cuda/archive/12.4.0/parallel-thread-execution/index.html.
  13. Andersch, M., Palmer, G., Krashinsky, R., Stam, N., Mehta, V., Brito, G., & Ramaswamy, S. (2022). NVIDIA Hopper Architecture In-Depth (NVIDIA Technical Blog). https://developer.nvidia.com/blog/nvidia-hopper-architecture-in-depth/.
  14. NVIDIA. (2023). Matrix Multiplication Background User’s Guide. https://docs.nvidia.com/deeplearning/performance/dl-performance-matrix-multiplication/index.html.
  15. NVIDIA. (2026). GPUDirect RDMA Documentation. https://docs.nvidia.com/cuda/gpudirect-rdma/.
  16. NVIDIA. (2016). Tesla P100 Datasheet. https://images.nvidia.com/content/tesla/pdf/nvidia-tesla-p100-datasheet.pdf.
  17. NVIDIA. (2018). Tesla V100 Datasheet. https://images.nvidia.com/content/technologies/volta/pdf/tesla-volta-v100-datasheet-letter-fnl-web.pdf.
  18. NVIDIA. (n.d.). NVIDIA A100 Tensor Core GPU specifications. https://www.nvidia.com/en-us/data-center/a100/.
  19. NVIDIA. (n.d.). NVIDIA H100 Tensor Core GPU product specifications. https://www.nvidia.com/en-us/data-center/h100/.
  20. NVIDIA. (n.d.). NVIDIA DGX B200 specifications. https://www.nvidia.com/en-us/data-center/dgx-b200/.
  21. NVIDIA. (n.d.). cuBLAS Library (cuBLAS 13.4 documentation). https://docs.nvidia.com/cuda/cublas/.
  22. Tillet, P., Kung, H. T., & Cox, D. (2019). Triton: An Intermediate Language and Compiler for Tiled Neural Network Computations. In Proceedings of MAPL 2019 (pp. 10-19). https://doi.org/10.1145/3315508.3329973.
  23. Triton project. (n.d.). triton.Config, Python API reference for the Triton main branch. https://triton-lang.org/main/python-api/generated/triton.Config.html.
  24. Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Re, C. (2022). FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness. https://arxiv.org/abs/2205.14135.
  25. Dao, T. (2023). FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning. https://arxiv.org/abs/2307.08691v1.
  26. Shah, J., Bikshandi, G., Zhang, Y., Thakkar, V., Ramani, P., & Dao, T. (2024). FlashAttention-3: Fast and Accurate Attention with Asynchrony and Low-Precision. https://arxiv.org/abs/2407.08608v2.
  27. Lefaudeux, B., Massa, F., et al. (2022). xFormers: A modular and hackable Transformer modelling library. https://github.com/facebookresearch/xformers.
  28. Chowdhery, A., et al. (2022). PaLM: Scaling Language Modeling with Pathways. arXiv:2204.02311.
  29. NVIDIA. (n.d.). nvidia-smi documentation. https://docs.nvidia.com/deploy/nvidia-smi/index.html.
  30. Williams, S., Waterman, A., & Patterson, D. (2009). Roofline: An Insightful Visual Performance Model for Multicore Architectures. Communications of the ACM, 52(4), 65-76. https://doi.org/10.1145/1498765.1498785.
  31. PyTorch. (2.14 documentation). torch.profiler. https://docs.pytorch.org/docs/2.14/profiler.html.
  32. NVIDIA. (n.d.). Nsight Systems User Guide. https://docs.nvidia.com/nsight-systems/UserGuide/index.html.
  33. NVIDIA. (n.d.). Nsight Compute Profiling Guide. https://docs.nvidia.com/nsight-compute/ProfilingGuide/index.html.
  34. PyTorch. (2.14 documentation). CUDA semantics. https://docs.pytorch.org/docs/2.14/notes/cuda.html.
  35. Ansel, J., et al. (2024). PyTorch 2: Faster Machine Learning Through Dynamic Python Bytecode Transformation and Graph Compilation. In Proceedings of ASPLOS 2024. https://doi.org/10.1145/3620665.3640366.
  36. PyTorch. (2.14 documentation). torch.compile. https://docs.pytorch.org/docs/2.14/generated/torch.compile.html.
  37. PyTorch. (2.14 documentation). RMSNorm. https://docs.pytorch.org/docs/2.14/generated/torch.nn.RMSNorm.html.
  38. Hu, E. J., et al. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv:2106.09685. https://arxiv.org/abs/2106.09685.
  39. PyTorch. (2.14 documentation). Automatic Mixed Precision examples, Gradient accumulation. https://docs.pytorch.org/docs/2.14/notes/amp_examples.html#gradient-accumulation.
  40. NVIDIA. (n.d.). Floating Point and IEEE 754, Sections 2.1-2.2 (CUDA 13.4 documentation). Retrieved September 29, 2026, from https://docs.nvidia.com/cuda/floating-point/index.html.
  41. NVIDIA. (Transformer Engine 2.19.0 documentation). Using FP8 and FP4 with Transformer Engine. https://docs.nvidia.com/deeplearning/transformer-engine/examples/fp8_primer.html.
  42. Micikevicius, P., et al. (2018). Mixed Precision Training. arXiv:1710.03740. https://arxiv.org/abs/1710.03740.
  43. Shazeer, N. (2020). GLU Variants Improve Transformer. arXiv:2002.05202. https://arxiv.org/abs/2002.05202.
  44. Meta Llama. (llama-models at commit 0e0b8c5, Oct 10, 2025). models/sku_list.py. https://github.com/meta-llama/llama-models/blob/0e0b8c519242d5833d8c11bffc1232b77ad7f301/models/sku_list.py.
  45. Meta Llama. (llama-models at commit 0e0b8c5, Oct 10, 2025). models/llama3_3/MODEL_CARD.md. https://github.com/meta-llama/llama-models/blob/0e0b8c519242d5833d8c11bffc1232b77ad7f301/models/llama3_3/MODEL_CARD.md.
  46. Shazeer, N. (2019). Fast Transformer Decoding: One Write-Head is All You Need. arXiv:1911.02150. https://arxiv.org/abs/1911.02150.
  47. Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate Post-Training Quantization for Generative Pre-trained Transformers. arXiv:2210.17323. https://arxiv.org/abs/2210.17323.
  48. Lin, J., et al. (2024). AWQ: Activation-aware Weight Quantization for On-Device LLM Compression and Acceleration. In Proceedings of MLSys 2024. arXiv:2306.00978. https://arxiv.org/abs/2306.00978.
  49. bitsandbytes-foundation. (n.d.). bitsandbytes. https://github.com/bitsandbytes-foundation/bitsandbytes.
  50. PyTorch. (torchao stable documentation). Welcome to the torchao Documentation. https://docs.pytorch.org/ao/stable/index.html.
  51. Subramanian, S., Saroufim, M., & Zhang, J. (2022, February 8). Practical Quantization in PyTorch. PyTorch. https://pytorch.org/blog/quantization-in-practice/.
  52. PyTorch. (2.14 documentation). MinMaxObserver. https://docs.pytorch.org/docs/2.14/generated/torch.ao.quantization.observer.MinMaxObserver.html.
  53. PyTorch. (2.14 documentation). torch.fake_quantize_per_tensor_affine. https://docs.pytorch.org/docs/2.14/generated/torch.fake_quantize_per_tensor_affine.html.
  54. Mishra, A., Albericio Latorre, J., Pool, J., Stosic, D., Stosic, D., Venkatesh, G., Yu, C., & Micikevicius, P. (2021). Accelerating Sparse Deep Neural Networks. arXiv:2104.08378.
  55. Frantar, E., & Alistarh, D. (2023). SparseGPT: Massive Language Models Can be Accurately Pruned in One-Shot. ICML 2023. arXiv:2301.00774.
  56. Sun, M., Liu, Z., Bair, A., & Kolter, J. Z. (2024). A Simple and Effective Pruning Approach for Large Language Models. ICLR 2024. arXiv:2306.11695.
  57. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the Knowledge in a Neural Network. arXiv:1503.02531.
  58. NVIDIA. (n.d.). L2 Cache Control (CUDA Programming Guide 13.4.2). Retrieved September 28, 2026, from https://docs.nvidia.com/cuda/cuda-programming-guide/04-special-topics/l2-cache-control.html.
  59. Kwon, W., et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. In SOSP 2023. arXiv:2309.06180. https://arxiv.org/abs/2309.06180.
  60. vLLM contributors. (n.d.). Performance and Tuning, Preemption (v0.6.4.post1 documentation). https://docs.vllm.ai/en/v0.6.4.post1/models/performance.html.
  61. Gu, A., & Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752. https://arxiv.org/abs/2312.00752.
  62. Jiang, A. Q., et al. (2023). Mistral 7B. arXiv:2310.06825.
  63. Xiao, G., Tian, Y., Chen, B., Han, S., & Lewis, M. (2024). Efficient Streaming Language Models with Attention Sinks. ICLR 2024. arXiv:2309.17453.
  64. DeepSeek-AI. (2024). DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model. arXiv:2405.04434.
  65. vLLM contributors. (n.d.). Engine Arguments (v0.6.4.post1 documentation). https://docs.vllm.ai/en/v0.6.4.post1/models/engine_args.html.
  66. Hugging Face. (n.d.). Utilities for generation. Retrieved September 28, 2026, from https://huggingface.co/docs/transformers/internal/generation_utils.
  67. Little, J. D. C. (1961). A Proof for the Queuing Formula: L = λW. Operations Research, 9(3), 383-387. https://doi.org/10.1287/opre.9.3.383.
  68. Agrawal, A., Kedia, N., Panwar, A., Mohan, J., Kwatra, N., Gulavani, B. S., Tumanov, A., & Ramjee, R. (2024). Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. https://arxiv.org/abs/2403.02310.
  69. Patel, P., Choukse, E., Zhang, C., Shah, A., Goiri, I., Maleki, S., & Bianchini, R. (2024). Splitwise: Efficient Generative LLM Inference Using Phase Splitting. https://arxiv.org/abs/2311.18677.
  70. Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., Barrett, C., & Sheng, Y. (2024). SGLang: Efficient Execution of Structured Language Model Programs. https://arxiv.org/abs/2312.07104.
  71. Liu, Y., Cheng, Y., Yao, J., An, Y., Chen, X., Feng, S., Huang, Y., Shen, S., Zhang, R., Du, K., & Jiang, J. (2025). LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM Inference. https://arxiv.org/abs/2510.09665.
  72. Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast Inference from Transformers via Speculative Decoding. https://arxiv.org/abs/2211.17192.
  73. Willard, B. T., & Louf, R. (2023). Efficient Guided Generation for Large Language Models. https://arxiv.org/abs/2307.09702.
  74. NVIDIA. (n.d.). Overview (TensorRT LLM 1.3.0rc28 documentation). Retrieved September 28, 2026, from https://nvidia.github.io/TensorRT-LLM/overview.html.
  75. NVIDIA. (n.d.). Migration Guide: TensorRT Backend Removed (TensorRT LLM 1.3.0rc28 documentation). Retrieved September 28, 2026, from https://nvidia.github.io/TensorRT-LLM/legacy/tensorrt-backend-removal.html.
  76. PyTorch. (2.14 documentation). torch.allclose. https://docs.pytorch.org/docs/2.14/generated/torch.allclose.html.
  77. PyTorch contributors. (2026). Common Graph Breaks (PyTorch 2.14 documentation). https://docs.pytorch.org/docs/2.14/user_guide/torch_compiler/compile/programming_model.common_graph_breaks.html.
  78. PyTorch. (2.14 documentation). torch.Tensor.copy_. https://docs.pytorch.org/docs/2.14/generated/torch.Tensor.copy_.html.
  79. Bindel, D. (2015, October 6). Distributed memory: Networks and models. Cornell University, CS 5220 course notes, Network properties. https://www.cs.cornell.edu/~bindel/class/cs5220-f15/slides/2015-10-06-network.html.
  80. NVIDIA. (2026). NVIDIA Collective Communications Library Documentation. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/index.html.
  81. NVIDIA. (2024). NCCL Tests PERFORMANCE Documentation. https://raw.githubusercontent.com/NVIDIA/nccl-tests/c6eb15875f508076f3f26de4f7da3899701bc4db/doc/PERFORMANCE.md.
  82. PyTorch. (2026). Distributed Communication Package: torch.distributed. https://docs.pytorch.org/docs/2.14/distributed.html.
  83. NVIDIA. (NCCL 2.32.3 documentation). Environment Variables. Retrieved September 28, 2026, from https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/env.html.
  84. PyTorch. (2026). Distributed Data Parallel. https://docs.pytorch.org/docs/2.14/notes/ddp.html.
  85. PyTorch. (2026). FullyShardedDataParallel. https://docs.pytorch.org/docs/2.14/fsdp.html.
  86. PyTorch. (2026). SGD. Versioned v2.14 API documentation at https://docs.pytorch.org/docs/2.14/generated/torch.optim.SGD.html.
  87. Ren, J., Rajbhandari, S., Aminabadi, R. Y., Ruwase, O., Yang, S., Zhang, M., Li, D., & He, Y. (2021). ZeRO-Offload: Democratizing Billion-Scale Model Training. USENIX ATC 2021. arXiv:2101.06840.
  88. Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., & Catanzaro, B. (2019). Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism. https://arxiv.org/abs/1909.08053.
  89. Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M. X., Chen, D., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., & Chen, Z. (2019). GPipe: Easy Scaling with Micro-Batch Pipeline Parallelism. https://arxiv.org/html/1811.06965v5.
  90. Lepikhin, D., et al. (2020). GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. https://arxiv.org/abs/2006.16668.
  91. Korthikanti, V., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., & Catanzaro, B. (2022). Reducing Activation Recomputation in Large Transformer Models. arXiv:2205.05198.
  92. Liu, H., Zaharia, M., & Abbeel, P. (2023). Ring Attention with Blockwise Transformers for Near-Infinite Context. arXiv:2310.01889.
  93. NVIDIA. (n.d.). Parallelisms Guide (Megatron Bridge documentation). Retrieved September 28, 2026, from https://docs.nvidia.com/nemo/megatron-bridge/0.4.1/parallelisms.html.
  94. Narayanan, D., et al. (2021). Efficient Large-Scale Language Model Training on GPU Clusters Using Megatron-LM. SC21. arXiv:2104.04473.
  95. Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. Journal of Machine Learning Research, 23(120), 1-39. arXiv:2101.03961.