Appendix D — Practice questions

These questions apply the book’s methods to design choices and failures. In each answer, identify the relevant component, show the calculation or records you would use, and explain what you would change or decide.

  1. Section 5.1: A support classifier reports 92% accuracy on a dataset with 94% non-escalation cases. Construct a stronger baseline and specify the confusion-matrix evidence needed before release.
  2. Section 3.1: A prompt contains policy, examples, a long conversation, retrieved evidence, tool schemas, and the user request. Show how you would allocate a strict context budget and record what was left out.
  3. Section 5.4: A judge prefers the longer of two otherwise equivalent answers. Design a pairwise evaluation that measures and reduces position and verbosity bias.
  4. Section 7.4: A FinanceBench-style answer is wrong even though the cited page contains the needed number. Name the records that distinguish a retrieval failure from context, generation, and deterministic-calculation failures.
  5. Section 6.6: Dense retrieval misses an exact product identifier while BM25 finds it, and dense retrieval finds a paraphrase that BM25 misses. Propose a hybrid and reranking experiment with offline metrics.
  6. Section 9.7: An agent succeeds in one demonstration but fails three of ten repeated trials. Define the task outcome, trajectory diagnostics, confidence reporting, and release implication.
  7. Section 10.3: A four-node plan has two independent middle tasks after a shared root and one final join. Derive its ready sets and explain why total work differs from critical-path latency.
  8. Section 10.6: An MCP server accepts a valid token intended for another resource server. Identify the authorization defect and the host/server evidence needed to prevent it.
  9. Section 11.1: A multi-agent design introduces a researcher, calculator, writer, and critic for a task one agent already completes reliably. State the evidence required to justify the split.
  10. Section 11.5: A user corrects a persistent preference and later asks for deletion. Explain how the system handles the conflicting records, saves the correction, retrieves it, and applies expiry and deletion rules to all copies, including indexes and caches.
  11. Section 7.5: A RAG improvement raises faithfulness but reduces page-hit at k and doubles p95 latency. Frame the release decision using slices and component responsibility.
  12. Section 8.8: Design the smallest sufficient architecture for a read-only assistant that resolves user questions over a private, frequently updated collection of policy documents with citations and an evidence-gap response.