Appendix E — Answer key
The answers below identify the key evidence and reasoning. Several designs may work if they address the question and include a way to test the choice.
E.1 Baselines and confusion-matrix evidence
Relevant teaching: Section 5.1, Metrics, output structure, and error cost.
The majority-class policy predicts non-escalation for every case. Because 94% of the dataset cases are non-escalations, that policy has 94% accuracy. The reported model accuracy is 92%, two percentage points below this baseline, so its accuracy does not justify release. Accuracy alone also cannot show whether the model finds the rare escalations that matter.
With requires escalation as the positive class, release evidence needs the actual-versus-predicted confusion matrix: TP for needed alerts, FP for unnecessary alerts, TN for correctly withheld alerts, and FN for missed escalations. Report precision and recall for each critical slice alongside the majority-class and current-production baselines. A model that exceeds 94% accuracy can still be unsafe if it misses a required escalation or fails another release gate.
E.2 Context-budget allocation
Relevant teaching: Section 3.1, Context-budget allocation.
With a shared input/output limit, the budget reserves space for output. Separate limits require separate checks. The input allocation protects system policy, current user intent, and the most task-relevant evidence. The measured value of examples, history, and tools determines their remaining allocation. The record retains the actual serialized request and relevant rendering settings alongside token counts, selection rules, truncation, summary versions and omitted IDs. Counts or selected IDs alone cannot reconstruct that request.
E.3 Pairwise judge calibration
Relevant teaching: Section 5.4, Judge-model calibration and robustness.
The calibration scores both A/B and B/A orderings, with order randomized across cases. To measure verbosity bias, meaning-equivalent short and long answers are compared in both orders. Selection rates by answer length are checked against audited human labels on held-out cases. Where the rubric permits, a length constraint reduces the opportunity for that bias. Raw rationales support error analysis, and order-sensitive disagreement remains visible instead of being averaged away.
E.4 Diagnosing a wrong cited answer
Relevant teaching: Section 7.4, Component-level and end-to-end failures.
The expected page identifies where the needed number should appear, but a page hit alone does not establish that the relevant passage was retrieved. The source file, parsed text, indexed chunks and retrieved IDs show whether preparation and search retained the number. The actual model request and truncation record then show whether it reached generation. If the retrieved passage contained the number but context assembly omitted it, that omission explains the missing input. If the request contained it but the answer misread it, generation requires investigation. Recomputing the arithmetic from the extracted values distinguishes a calculation error from an incorrect input value. Several failures can coexist.
E.5 Hybrid retrieval and reranking
Relevant teaching: Section 6.6, Staged retrieval pipeline.
A hybrid experiment retrieves candidates from both indexes, removes duplicate source units and compares rank fusion with each retriever alone on shared held-out questions. A separate candidate adds a reranker so its contribution can be measured. Page-hit rate, retrieval precision and recall, MRR, latency, and identifier and paraphrase slices use the same labels and cutoffs across candidates. A gain on one question does not establish an aggregate improvement.
E.6 Repeated-trial agent reliability
Relevant teaching: Section 9.7, Agent outcomes across repeated trials.
Success is defined by the final environment or validated answer state. Independent trials yield an observed success rate and an interval. Tool, argument, error, and stopping records explain the failures. Seven successes in ten trials gives an observed rate of 0.70, with a 95 percent Wilson interval of about 0.40 to 0.89. Whether that blocks release depends on the task’s required reliability and the uncertainty in the estimate. One successful demonstration does not answer that question.
E.7 Ready sets and critical-path latency
Relevant teaching: Section 10.3, Dependency-aware parallel scheduling.
Initially only the root is ready. After it completes, both independent middle nodes are ready, and the join becomes ready only after both complete. Total work sums all durations, while lower-bound latency follows the root plus the slower middle branch plus the join.
E.8 MCP token audience validation
Relevant teaching: Section 10.6, Trust at the protocol boundary.
The server accepted a token intended for a different resource and must reject it. For the OAuth-based MCP setup taught in Section 10.6, the host’s client requests a token for that resource, and the server validates the audience and other required token claims before enforcing access to the requested operation. Audience binding reduces the risk of misusing authority across services but does not prevent every confused-deputy attack. The server must not pass a token issued to it unchanged to a downstream service. The test records the intended resource, validation result and rejected operation without exposing the credential.
E.9 Evidence for adding agents
Relevant teaching: Section 11.1, When extra agents are warranted.
The case for extra agents starts with a task requirement or a measured single-agent limitation, such as unavailable tool specialization, context overflow, a need to isolate responsibilities, or an opportunity to shorten the critical path. Handoff loss, coordination cost, new failure modes, and reliability are then compared with the single-agent baseline.
E.10 Memory correction and deletion
Relevant teaching: Section 11.5, Semantic memory and correction.
The application checks whether the new statement corrects the same fact for the same identity and purpose. An accepted correction supersedes the old record so retrieval does not present both as current. Identity, access and expiry checks run before ranking eligible records by relevance, freshness and source reliability.
A deletion request first makes the record ineligible for retrieval, then propagates removal to live indexes, caches and replicas. Cleanup failures remain pending and are retried rather than reported as complete. Tests check that stale replicas cannot return or restore the value. Any retained backups require a separate retention and restore policy so restoring a backup does not reactivate the deleted record. The audit stores only the identifiers and status needed to verify this process, not the removed value.
E.11 Trade-offs across RAG metrics
Relevant teaching: Section 7.5, One-change improvement cycles.
The apparent gain must be read alongside its losses in other measures. Missed labeled evidence calls for retrieval analysis, while a slice-level review checks faithfulness. Correctness, latency, and cost are compared with their thresholds. Release depends on all critical gates passing.
E.12 Smallest sufficient RAG architecture
Relevant teaching: Section 8.8, The smallest sufficient workflow.
Ingestion retains each document’s source, version, and permissions. Source changes and deletions update or remove the corresponding indexed chunks.
At request time, the current access rules apply again. The retrieval method starts with the simplest approach that meets the measured evidence requirements. Hybrid retrieval becomes useful if tests reveal complementary exact-term and paraphrase failures. Recorded chunking and retrieval settings make the selection reproducible, while fixed code assembles the passages. The answer carries structured citations, and an evidence-gap response handles cases where the sources do not support an answer.
Retrieval, support for the answer, and latency are evaluated separately. This task can use a fixed workflow without an agent choosing the next action.
Assessment standard
A useful answer explains where the failure arose, which records or measurements support that explanation, and what decision follows. It also identifies uncertainty that the available evidence cannot resolve.