Appendix B — Questions and concise answers
A production failure can cross artifacts, compute, storage, serving, retrieval, evidence, and orchestration. The questions follow the chapters in order. Conceptual questions ask which component, evidence, or recovery path explains a symptom or decision. Calculation questions reuse the book’s formulas with new numbers so that the reasoning, not the example values, is tested.
B.1 Questions
Each question begins with a system symptom, a decision, or a set of measurements. Questions marked Calculation have a numeric answer followed by its interpretation.
Chapter 1: Reproducible runs
- Which versioned information is needed to reconstruct how a model serving result was produced?
- Calculation. Peak demand is 90 requests/s. One replica sustains 10 requests/s at the latency target, the plan allows 80% utilization, and two spare replicas cover a host-group loss. How many replicas does the starting deployment need?
- Why does an image digest not fully identify the runtime that a GPU training job sees?
Chapter 2: Distributed training
- When is DDP an inadequate parallelism strategy even if communication is fast?
- Why does gang scheduling matter for a multi-node distributed job?
- How do RPO and RTO affect checkpoint frequency, placement, and restore testing?
- Calculation. A 13-billion-parameter model uses the same mixed-precision Adam layout as Section 2.1 (16 bytes per parameter). What is the persistent state, and how much remains per GPU after ideal eight-way full sharding?
- Calculation. Sixteen ranks run a ring all-reduce on a 2 GB gradient bucket. How many bytes does each rank send, and how long does that take at an effective 50 GB/s per rank if nothing overlaps it?
- Calculation. A pipeline has 8 stages. What is the bubble fraction with 8 micro-batches and with 32? What does activation recomputation add to the arithmetic of each step?
- Calculation. A 64-GPU run processes 400,000 tokens/s with 60 GFLOPs of useful model arithmetic per token. Each GPU has a 1.0 PFLOP/s dense peak at the run precision. What is the MFU?
Chapter 3: Storage for training
- How do block, shared-file, and object storage semantics differ for ML artifacts?
- What makes a distributed checkpoint generation safe to publish to restore clients?
- Why can high storage IOPS coexist with low bandwidth?
- Calculation. A storage path completes 50,000 read operations per second. What is its throughput at 256 KiB per operation, and at 16 KiB?
- Calculation. A job saves every 20 minutes of useful work, each synchronous save takes 2 minutes, and restore after a node loss takes 5 minutes. Assuming failures are uniform over useful work, what is the expected interruption per failure, and what share of each cycle is save time?
Chapter 4: Predictable serving
- Why are TTFT and ITL separate serving objectives?
- How does continuous batching improve throughput, and what latency risk does it introduce?
- Why does KV-cache state influence routing and admission control?
- Calculation. A model has 40 layers and 8 key/value heads of width 128 with a BF16 cache. After weights and buffers, 30 GB remain for KV blocks. How many bytes does one cached token use, and how many concurrent 4,000-token sequences fit?
- Calculation. A request has a TTFT of 500 ms and generates 250 tokens at an average ITL of 20 ms. How long does it take? A two-GPU replica priced at $4 per GPU-hour sustains 3,000 output tokens/s. What is the cost per million output tokens, and how many 250-token answers per second does it serve?
- Calculation. A 34-billion-parameter model is served on 80 GB GPUs used at 0.90, with 4 GB of buffers per GPU. How much KV memory remains with BF16 weights on one GPU, with 8-bit weights on one GPU, and with BF16 weights on two tensor-parallel GPUs?
- Calculation. Speculative decoding proposes 3 tokens per step. Assume each proposal is accepted independently with the same probability, as in the simplified calculation in Section 4.6. How many tokens does one target pass yield on average at an acceptance probability of 0.8 and of 0.4, and when is the technique worth using?
Chapter 5: Retrieval for grounded answers
- Why is an embedding-model revision a schema migration for a vector system?
- How do HNSW and IVF-PQ differ in memory use, search, and tuning behavior?
- Why does hybrid retrieval use rank fusion or calibrated scores instead of arbitrary raw-score addition?
- Which measurements belong in an honest vector-system benchmark?
- Calculation. What is the cosine similarity of vectors (3, 4) and (4, 3)? With \(k = 60\), chunk A is ranked 1 by dense search and absent from sparse results, while chunk B is ranked 4 by dense and 2 by sparse. Which ranks first after RRF?
- Calculation. A question has three relevant chunks. The top five results have relevance 1, 0, 1, 0, 0. What are recall@5, precision@5, the reciprocal rank, and binary nDCG@5?
Chapter 6: Measurement and evaluation
- Why can an HTTP 200 response coexist with a serious ML-system failure?
- What do logs, metrics, and traces each contribute to a quality incident?
- Why are both correctness and consistency needed for an LLM judge?
- How can a prompt decision tree expose evaluator defects?
- Why do AI traces preserve retrieved source identities while applying content governance?
- Calculation. On 100 adjudicated answers, a hallucination detector has TP = 18, FP = 6, FN = 2, and TN = 74. What are its accuracy, precision, recall, F1, and Cohen’s kappa against the reference?
Chapter 7: Repeatable workflows
- What is the difference between orchestration metadata and experiment-tracking metadata?
- What makes a pipeline task idempotent across retries and backfills?
- Why is production-alias promotion separated from a compute task?
- Calculation. Validation takes 5 minutes. One branch then embeds (30), builds the index (15), validates it (5), and runs the evaluation suite (20). Another branch calibrates the judge (12). A 4-minute promotion gate waits for both. What is the lower bound on run time, and which change would shorten it?
Chapter 8: End-to-end operation
- In the Chapter 8 incident, which evidence made an inference-capacity problem less likely, and what pointed toward retrieval?
- Where does the prevention control for an incomplete vector collection belong?
- Calculation. The TTFT objective is p95 \(\leq\) 1,000 ms. A team uses per-stage p95 observations of gateway 25 ms, query embedding 20 ms, search 70 ms, reranking 150 ms, packing 5 ms, queue 120 ms, and prefill 500 ms as a chosen deterministic planning allocation. Can a 150 ms query-rewriting model call fit that allocation, and what evidence is needed before claiming that the end-to-end p95 objective is met?
B.2 Concise answers
Each answer identifies the relevant system requirement and evidence. Calculation answers show the arithmetic and then state what the result means for the decision.
- The serving bundle records the weights, tokenizer, engine image, precision or quantization, adapters, prompt configuration, and hardware compatibility. A grounded-answer release also records the retriever configuration, collection, source corpus, and shared release identity. These records allow the request to be reconstructed and a new response compared with the saved answer. They do not guarantee identical generated text.
- Each replica is planned at 10 × 0.8 = 8 requests/s, so ⌈90 ÷ 8⌉ = ⌈11.25⌉ = 12 demand replicas, plus 2 spares = 14. A load test with the real prompt and output mix confirms the estimate.
- The digest fixes the image’s files and configuration, but the host driver, kernel, and GPU hardware come from the node. Reproducing a run also requires those host facts to be recorded and tested with the image. Numerical results may still differ across platforms or software releases.
- DDP replicates the full model and optimizer-related state. If that state does not fit on one device, state sharding can reduce the per-device copy. Tensor, pipeline, context, or expert parallelism address different model, sequence, or architecture limits. Their communication and placement costs must be measured. Global-batch or data constraints alone do not select one of those strategies.
- The scheduler admits the configured group together, then the ranks start and rendezvous. Partial allocation can waste reserved GPUs while the group remains incomplete. Group admission does not ensure simultaneous readiness or prevent later failure.
- RPO bounds acceptable lost work and helps choose the interval, while RTO constrains where checkpoints live and how fast they restore. A policy is unproven until a fresh allocation restores it.
- 13 × 10⁹ × 16 bytes = 208 GB of persistent state. Eight-way full sharding leaves 208 ÷ 8 = 26 GB per GPU before activations, temporary all-gathers, and allocator reserve. On an 80 GB GPU the remaining room goes to those temporary costs, so the run may fit where DDP could not.
- 2 × (16 − 1) ÷ 16 × 2 GB = 3.75 GB sent per rank, with the same amount received. At 50 GB/s that is 3.75 ÷ 50 = 75 ms per bucket if no computation overlaps it, which is why DDP overlaps buckets with the backward pass.
- With 8 micro-batches, (8 − 1) ÷ (8 + 8 − 1) = 7 ÷ 15 \(\approx\) 47% of the schedule is idle. With 32, 7 ÷ 39 \(\approx\) 18%. Full activation recomputation adds about one forward pass, raising step arithmetic by about one third. More micro-batches cut the bubble but make each kernel smaller.
- Observed rate = 400,000 × 60 × 10⁹ = 24 PFLOP/s. Peak = 64 × 1.0 = 64 PFLOP/s. MFU = 24 ÷ 64 = 37.5%. The remaining 62.5% needs timelines to attribute it to input waits, communication, bubbles, or kernel efficiency.
- Block exposes a disk-like device, shared file exposes a multi-client hierarchical namespace, and object storage exposes keyed objects through an API. Their update, rename, sharing, and durability behavior differ.
- Expected shards and compatible state are durably stored across the required failure domains and restore-tested before a coordinator conditionally publishes the immutable approved generation. Stale writers are rejected, and readers ignore incomplete generations. This is not a multi-shard transaction supplied by a marker alone.1
- Bandwidth equals IOPS times bytes per operation. Small operations can yield many IOPS but few transferred bytes.
- At 256 KiB: 50,000 × 256 KiB = 12,500 MiB/s, about 12.2 GiB/s or 13.1 GB/s, if the network and servers can sustain it. At 16 KiB: 50,000 × 16 KiB \(\approx\) 781 MiB/s. The same IOPS figure describes a sixteenfold difference in bandwidth.
- Expected lost work is half the interval, 10 minutes, and restore adds 5, for 15 minutes per failure. The save uses 2 ÷ 22 \(\approx\) 9.1% of each cycle. A shorter interval would cut lost work but raise that share.
- TTFT covers queue plus prompt processing before the first token, while ITL describes decode responsiveness. Different optimizations and user perceptions apply to each stage.
- Continuous batching fills freed batch slots and increases utilization under variable lengths. Aggressive admission can increase queueing, memory pressure, or ITL for active requests.
- The cache consumes capacity proportional to active context and may be reusable for shared prefixes. Routing considers capacity and locality, while admission prevents overcommit.
- Per token: 2 × 40 × 1,024 × 2 = 163,840 bytes. 30 × 10⁹ ÷ 163,840 \(\approx\) 183,000 tokens, or about 45 sequences of 4,000 tokens. Allocator metadata and partly filled blocks reduce the real figure slightly.
- Latency: 500 + 249 × 20 = 5,480 ms. Cost: $8 per hour × 10⁶ ÷ (3,600 × 3,000) ≈ $0.74 per million output tokens. Request rate: 3,000 ÷ 250 = 12 answers/s. This calculation assumes a fixed sustained output-token rate: under that assumption, longer answers lower the request rate and raise cost per request while per-token compute cost is unchanged. A changed prompt or output mix needs a new benchmark.
- BF16 weights take 68 GB. One GPU: 72 − 68 − 4 = 0 GB, so the placement fails. 8-bit weights: 72 − 34 − 4 = 34 GB, if the quantized model passes its quality gates. Two tensor-parallel GPUs: 144 − 68 − 8 = 68 GB, at the cost of a second GPU and per-layer all-reduce.
- At α = 0.8: (1 − 0.8⁴) ÷ (1 − 0.8) \(\approx\) 2.95 tokens per target pass. At α = 0.4: (1 − 0.4⁴) ÷ 0.6 \(\approx\) 1.62. The technique helps latency-sensitive traffic when acceptance is high and the draft model is cheap. At low acceptance or large, compute-bound batches, the draft work can cancel the gain.
- An incompatible model change needs a new vector configuration and a matching query encoder. Backfill catches up concurrent updates, deletions, and ACL changes to one watermark. Source changes remain paused through cutover, or dual writes and replay continue until the new pair serves traffic. Evaluation precedes a coordinated descriptor change, each request uses one encoder-collection pair, and rollback must still meet current access and deletion rules.2
- HNSW traverses a layered proximity graph and tunes a configured working-set bound. IVF-PQ probes selected clusters and scores compressed codes. Recall, latency, memory, filtering, and update costs must be measured for the actual corpus and settings.
- Dense, sparse, and reranker raw scores have different scales and directions. Rank fusion uses rank positions instead of arbitrary raw-score addition. Score calibration is a separate method and is not performed by RRF.3
- An honest benchmark reports recall, precision, ranking quality, p50/p95/p99 latency, resource and index size, concurrency and filters, build/update/delete freshness, failure and recovery, and end-to-end groundedness.
- Cosine: 24 ÷ (5 × 5) = 0.96. RRF: A = 1/61 \(\approx\) 0.0164, and B = 1/64 + 1/62 \(\approx\) 0.0318, so B ranks first because both retrievers placed it well.
- Recall@5 = 2 ÷ 3 \(\approx\) 0.67, precision@5 = 2 ÷ 5 = 0.40, and reciprocal rank = 1 because the first result is relevant. DCG = 1 + 1/log₂4 = 1.5, and the ideal ordering gives IDCG = 1 + 1/log₂3 + 1/log₂4 \(\approx\) 2.131, so nDCG@5 \(\approx\) 0.70. The first position is right, but a third of the relevant evidence is missing.
- Transport success proves only that a request returned. The model can still be wrong, unsafe, stale, ungrounded, or silently degraded. Quality evidence and business outcomes are separate from availability.
- Metrics reveal population change, traces show the timed request path, and logs preserve individual events. Correlation IDs join them. A trace helps locate a failure but does not by itself prove its cause.
- A judge can agree repeatedly and still be wrong. Correctness compares its verdict with a reference. Consistency measures agreement across repeated judgments.
- It makes precedence, overrides, missing branches, contradictions, and ambiguous scope explicit while requiring each node to anchor to prompt text.
- Source identities support replay, citations, and failure localization. Full prompts or documents may contain sensitive data, so policy governs sampling, redaction, encryption, access, and retention.
- Accuracy = (18 + 74) ÷ 100 = 0.92. Precision = 18 ÷ 24 = 0.75, recall = 18 ÷ 20 = 0.90, F1 = 2 × 0.75 × 0.90 ÷ 1.65 \(\approx\) 0.82. The reference is 20% positive and the detector 24% positive, so \(p_{\mathrm{e}} = 0.20 \times 0.24 + 0.80 \times 0.76 = 0.656\) and κ = (0.92 − 0.656) ÷ (1 − 0.656) \(\approx\) 0.77. The detector catches most hallucinations, while one in four flags is a false alarm.
- The orchestrator owns DAG, task, schedule, and retry state. The experiment tracker owns comparable parameters, metrics, model and evaluator versions, and research runs. Stable IDs link them.
- It binds work to immutable inputs and a logical interval, writes attempt-scoped output, validates, and commits one generation so repetition does not duplicate or corrupt effects.
- Retries could repeat the side effect or expose partial state. A separate gated promotion task updates the alias after validation and can enforce approval or rollback authority.
- The retrieval path takes 5 + 30 + 15 + 5 + 20 + 4 = 79 minutes and the judge path 5 + 12 + 4 = 21, so 79 minutes is the lower bound. Speeding the judge calibration changes nothing. Halving embedding to 15 minutes lowers the bound to 64 minutes.
- Stable queue, TTFT, ITL, memory, throughput, and HTTP error signals make an inference-capacity fault less likely, but do not rule it out for every request. Failures cluster in the new-corpus slice, and traces lack expected source identities. Those observations point toward retrieval coverage and the collection build.
- Before publication, expected partitions, counts, and hashes are reconciled, and the promotion task changes the production alias only after a validation manifest is committed.
- The chosen allocation totals 25 + 20 + 70 + 150 + 5 + 120 + 500 = 890 ms. Adding 150 ms gives 1,040 ms, so the call does not fit that allocation unless another allocation is reduced, for example by reranking fewer candidates. This arithmetic is not the end-to-end p95: component p95 values alone do not determine the p95 of their sum. Measure end-to-end TTFT under the expected workload, or validate a joint latency model, before claiming that the p95 objective is met.
Amazon Web Services. (n.d.). What is Amazon S3? Amazon Simple Storage Service User Guide. https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html#ConsistencyModel. S3 supplies atomic single-key updates, not a checkpoint transaction. Sections 3.4 and 7.6 add the complete conditional protocol.↩︎
Qdrant. (n.d.). Migrate to a new embedding model with zero downtime in Qdrant (Operations documentation). https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/. Qdrant supports side-by-side migration and dual writes. Coordinated query-encoder release and current-policy rollback are additional book requirements.↩︎
Cormack, G. V., Clarke, C. L. A., & Büttcher, S. (2009). Reciprocal rank fusion outperforms Condorcet and individual rank learning methods. In Proceedings of SIGIR 2009 (pp. 758–759). ACM. DOI: 10.1145/1571941.1572114. https://cormack.uwaterloo.ca/cormacksigir09-rrf.pdf. The original RRF formula combines reciprocal rank positions rather than calibrated raw scores.↩︎