Part V: Adapting and Running LLMs

A pretrained model may lack relevant information, fail a required format, or generate too slowly for a task. Changing context, changing parameters, and scheduling inference affect different objects. Each choice needs evidence against the intended task and resource limits.

Chapter 20 compares prompting, retrieved evidence, full fine-tuning, and smaller trainable additions. It separates the objective guiding an update from the storage format of the frozen base. Quantization can reduce weight payload while changing the numerical values used in computation.

Chapter 21 follows changing token and attention state with fixed model parameters in its cache and scheduling traces. It also compares query-head arrangements, which are model architecture choices. Its cache timeline supports memory estimates, request latency, and block allocation. Part VI then connects other input representations, tensor storage, and experiment evidence.

Two chapter panels show adapter and cache paths. The upper adapter path omits the parallel sum, and prefill lacks its first-token selection output.
Figure V.1: The panels connect adaptation choices with prompt processing, decoding, and stored attention state.

Chapters in this part