13 Evidence for release decisions
A candidate can give the right answer in a test and still use unauthorized data, repeat a tool call, or miss its latency target. A release review checks the complete result and the records of how the system produced it. The decision applies to one system version, its permitted inputs and actions, and the conditions tested.
The map connects architecture choices, failure diagnosis, a release record, and monitoring after release. Each stage uses evidence from the model, retrieval, tools, and memory only where the selected system needs them.
13.1 Smallest sufficient architecture
The examples in Chapter 12 combine different components for different tasks. Adding all of them to a new product would add costs and failure paths without necessarily improving the result. The choices below start from the task and its tests, then add capabilities for requirements the current design cannot meet. This is the book’s engineering selection rule, not a fixed sequence that every product must follow.
- The task specification states the user-visible behavior, allowed inputs, required final state, non-goals, safety constraints, latency target, and cost target.
- Representative and adversarial examples, a baseline, and versioned evaluation records establish evidence before a design change.
- One model call is a baseline when the required information fits the request and no external action is needed. It is sufficient only if its measured results meet the task’s quality, safety, latency, and cost requirements.
- Deterministic context assembly supplies policy, history, examples, or structured records to the call without giving the model control over their selection.
- Retrieval supplies information from allowed external sources. Document search and structured queries solve different access problems, as Chapter 6 explains. Retrieval and generation remain separate evaluation targets.
- A workflow keeps known steps and branches in software, including branches selected from observations.
- An agent lets the model choose the next permitted tool action from observations when the useful path is not known in advance. Application code still limits actions and steps.
- Explicit planning makes dependencies and remaining work visible. MCP shares capability discovery and call formats. Multiple agents divide work among specialized loops. Long-term memory preserves selected information across sessions. Each is a separate choice justified by the task and its tests.
For example, a fixed workflow can query a database and validate the result without document retrieval or an agent. Adding MCP changes how an allowed tool is discovered and called, not whether the model needs to choose the next tool. The table relates each observed need to a likely design and the evidence needed to justify it.
| Observed need | Smallest likely pattern | Evidence before escalation |
|---|---|---|
| Stable-rubric transformation of supplied text | Prompted model call with structured output | Golden set, schema validity, quality, latency, cost |
| Private or current document access | RAG pipeline | Naive baseline, retrieval labels, support and citation checks |
| Routing among known paths | Deterministic or model-assisted workflow | Route confusion matrix and branch tests |
| Iterative tool choice | ReAct agent with a step limit | Tool-selection tests, max-step fallback, repeated final-state success |
| Action dependencies need a revisable plan | Plan-and-Act | Plan validity, replan value, latency and cost delta |
| Independent expensive subtasks | Dependency graph and parallel scheduler | Dependency audit and critical-path measurement |
| Shared capabilities across hosts or teams | MCP interface | Measured integration burden, permission and transport tests |
| One agent hits tool, context, serial, or domain responsibility limits | Orchestrator-worker or justified alternative | Single-agent baseline and a metric for the measured limitation |
| Useful state must survive sessions | Typed episodic or semantic memory | Retrieval utility, correction, retention, and privacy tests |
13.2 Failure localization before model changes
The selected architecture identifies the components that can affect a result. When an evaluation fails, replacing the model before identifying the cause can add unrelated output changes. A symptom narrows the investigation but does not prove which component failed. This table gives possible causes, records to inspect, and a controlled diagnostic test. If a fact was lost, the investigation follows it from the source through ingestion, retrieval, and the actual model request before judging generation.
| Symptom | Possible causes | First inspection | Controlled next test |
|---|---|---|---|
| Answer ignores a source fact | Ingestion, retrieval, context assembly, or generation | Is the fact preserved in parsed text, indexed chunks, retrieval results, and the actual model request? | Compare source, parsed text, index entry, retrieval result, and request snapshot before testing generation |
| Correct evidence, unsupported claim | Generation or validation | Does the claim follow from the supplied chunks? | Faithfulness judge plus deterministic numeric check |
| Agent calls a plausible wrong tool | Tool schema, router, model selection, or validation | Are names, descriptions, arguments, and route classes distinct? | Tool-choice golden set and description ablation |
| Agent loops | Stopping policy or tool feedback | Is progress represented in state and is the fallback reachable? | Forced no-progress trace with step budget |
| Follow-up loses reference | Workflow checkpoint or context assembly | Was session state restored under the same ID? | Restart and pronoun-resolution test |
| Old preference overrides correction | Semantic-memory conflict policy | Were both facts appended instead of superseded? | Correction, precedence, and deletion test |
| Multi-agent output is incomplete | Handoff, worker output, or join logic | Did every required worker return a valid output, and did the join retain it? | Missing-worker, malformed-output, partial-failure, and join-omission tests |
| MCP tool writes too broadly | Authorization and server scope | Which approved user intent, token permissions, and server access policy allowed the write? | Denied-path, expired-token, and approval test |
One-variable discipline
A controlled diagnostic changes one suspected component and states the expected metric change. It keeps the test cases and other relevant settings fixed, records the changed version, and checks effects on quality, safety, latency, and cost. A model change is justified when that investigation points to model behavior rather than a missing input or a failed control.
13.3 The release decision record
Section 13.2 established how to localize a failed test case. After an evaluation run, an accountable reviewer must decide whether the measured results and enforced controls support release.
- Release decision record: A versioned document that links the task definition, evaluation results, component traces, thresholds, blocking failures, and accountable follow-up to a PASS or HOLD decision for specified operating conditions.
Reviewers use the release decision record to compare tested behavior with release requirements. A high aggregate score can conceal local failures: a model can produce a correct answer from unauthorized documents, or an agent can succeed inconsistently across repeated trials. Success on the last allowed attempt is still success if it meets the outcome, safety, latency, and cost requirements. The record places task, retrieval, control, and infrastructure evidence beside their thresholds. A change that alters tested behavior or risk requires a new record. Examples include adding a write-capable tool, broadening permissions, connecting a new MCP server, adding an agent route, storing a new memory category, or changing a safety threshold.
The record uses the evaluation suite introduced in Chapter 4 and execution traces like those in Chapter 12. If the golden set omits edge cases, reviewers may overestimate readiness. The following illustrative record extends the support-data example from Chapter 12 with release requirements and test results. Those results are assumed for the example, not measurements from a provider run or controls implemented by the shortened Chapter 12 listing.
A compact record contains system_version, task_scope, thresholds and slice results, source or index versions, tool permissions and budgets, protocol and memory versions when used, p95 latency, cost per case, the PASS or HOLD decision, blocking failures, the person accountable for follow-up, and the required work. The four sections below identify the evidence needed for those fields in the illustrative refund-count release.
Example: A release record leads to HOLD
The following illustrative record uses the refund-count trace from Sections 9.8 and 12.5. The tool count and displayed count are both 27, so that check passes. Intent-routing edge cases pass in 93 of 100 cases, below the required 95 of 100. Unauthorized tool attempts are rejected in 20 of 20 cases. The p95 latency is 760 ms against an 800 ms limit. Cost is $0.018 against a $0.020 limit.
Decision: HOLD. The routing slice fails a required threshold despite the passing count, control, latency, and cost checks. Repair the routing errors using development cases, then rerun the versioned validation suite. If the repair used failed validation cases, a fresh held-out sample is needed to check generalization. Twenty rejected unauthorized attempts show the result for those tests, not a guarantee for all possible calls. MCP, multiple agents, and persistent user profiles are not used in this example, so the record marks them not applicable.
13.3.1 Task and evaluation evidence
The release review begins with the evaluation lineage established in Chapter 4. A raw accuracy score does not explain what the system is allowed to do or whether it handles edge cases safely. The first section of the release record therefore defines success and summarizes the metric outcomes.
The task specification should list the final state, non-goals, and automatic failures. For the refund-count agent, the required final state is a count of permitted rows that matches the recorded dataset version and intent filter. Policy questions and technical-support requests are outside this route and need another permitted route or an out-of-scope response. The record includes baseline metrics, candidate results, and any judge calibration used for answers written in natural language. It breaks performance down by critical slices. One field fragment is task_scope: refund_count, intent_routing: 93/100, and required_intent_routing: 95/100. Separate results for normal, adversarial, and empty-input slices show which cases caused the shortfall. The sample size and threshold belong to this example’s release policy, not a universal readiness standard.
13.3.2 Retrieval and grounding evidence
Task evaluation checks the final answer against the requirements. The record should also show whether the system used authorized information of an appropriate version. The evidence path depends on the route: document retrieval follows indexed chunks through the model request to the answer and its citations, while a structured-data route follows access checks, filters, and tool results. Appendix A lists trace fields and operational evidence records.
An answer can look correct even when its trace omits source identity or freshness, or when application code read data outside the caller’s permissions. The release record traces each route’s information path. For the refund-count route, an illustrative field fragment records dataset_version: support-2026-09-01, caller_scope: approved_support_dataset, intent_filter: get_refund, row_ref: result-4821, and count: 27. Application code checks the caller’s permitted dataset before applying the intent filter. The value get_refund selects a category, not a permission. No document retrieval or citation is claimed. A policy-question route needs separate tests for ingestion and index freshness, the chunks preserved in the model request, supported citations, and access control.
13.3.3 Agent control evidence
The evidence path checks the information used. The control record must also show that application code permits only authorized actions and stops the loop within its configured budget. It documents state schemas, registered tools, permission checks, and execution limits.
Dynamic tool choice introduces execution risks that answer-only evaluation cannot expose. The release record lists every tool registered for the agent loop, including its schema, access checks, and error behavior. For the illustrative refund-count release, the field fragment is tools: [filter_by_intent, count_rows], dataset_write_tools: [], step_budget: 3, and unauthorized_attempts_rejected: 20/20. The filter tool still stores a temporary result record, as Section 9.2 explains. That record needs caller access checks and a retention rule even though the support dataset is unchanged. One step must have a defined meaning. In Section 12.5 it is a routing or tool attempt, so the successful path spends one step on routing and two on tools. A full controller also needs failure tests showing that invalid proposals consume the budget, that no further call executes at exhaustion, and that the caller receives a failure or escalation status. Unauthorized-caller and stale-version tests are separate from that stopping test. This is evidence required from the implementation, not behavior supplied by the release record itself.
13.3.4 Infrastructure evidence
Agent controls limit an individual loop. Shared servers, multiple agents, and persistent memory add controls that one loop’s tests cannot establish. The infrastructure section records the interfaces, access policies, and storage rules for the components actually used.
A controlled agent can operate without MCP, multiple agents, or long-term memory. The refund-count example records mcp: not_applicable, persistent_memory: not_applicable, p95_latency_ms: 760, and cost_per_case: 0.018. Latency needs a workload and measurement window, and cost needs a currency and a rule for counting model and tool charges. Here the cost is US dollars per evaluated case. If the system later uses MCP to access a billing server, the record identifies the protocol and server versions, the credential identity and permissions, and the allowed capabilities. It must not copy secret tokens into the record. Chapter 10 explains the version-dependent protocol details and the access checks that remain the host’s and server’s responsibility. If the system stores user preferences, deletion tests cover the primary store, derived indexes, caches, and any other retained copies required by its policy. Retrieved memory is data, not authority to change policy or execute a tool. Appendix D poses practice cases for diagnosis and release decisions, and Appendix E answers them. The References section lists the cited sources.
13.4 Operating rule for release decisions
An added component should meet a task requirement or address a measured limitation, with enough benefit to justify its costs and failure paths. Mandatory safety and access controls remain release conditions, not costs that an accuracy gain can outweigh. The decision record keeps task results beside information sources, tool and state traces, safety checks, latency, and cost. If a critical slice or required control fails, the result is HOLD even when the average score improves. A PASS supports only the tested version and operating conditions, with the uncertainty of the tests used. Later changes require a new comparison.
Passing offline tests does not establish a benefit on live traffic. If the candidate improves an offline metric, the product outcome identified in Section 4.1 is still a hypothesis to check after release. The NIST AI RMF lists post-deployment monitoring, user input, appeal and override, incident response, and change management among its outcomes (MANAGE 4.1).1 The following methods support different parts of that check:
Shadow deployment: The candidate processes copies of live requests while users receive only the current system’s answers. Comparing both outputs tests the candidate on real traffic. Tool calls with side effects are blocked or simulated.
Canary release: A partial and time-limited deployment of a change, evaluated before full rollout.2 Its stop rules are written before it starts.
Online controlled experiment: Random assignment of users or sessions to the current and candidate versions, followed by a comparison of product outcome metrics.3 It estimates the change in the chosen outcome under the experiment’s assignment and measurement conditions. Sample size, uncertain estimates, and interactions between groups can limit the conclusion.
Service level objective (SLO): A target value or range for a measured service indicator, such as p95 latency or error rate.4 The service policy defines the measurement window and the response to a breach, such as pausing rollout or rolling back.
Shadow deployment compares responses without exposing users to candidate decisions. A canary limits live exposure while monitoring failures. A controlled experiment estimates a product-outcome difference between assigned groups. These approaches can be combined, but none replaces access checks or the offline release requirements. A canary may also use randomized assignment. Monitoring reliability is a different question from measuring a product benefit.
After PASS, the refund-count agent could serve a selected 5 percent of analysts for one week while the remaining analysts use the current version. Before rollout, the record defines how analysts are assigned, the observation window and minimum request count for p95 latency, and who can stop the rollout. The platform measures the canary and current version separately. In this example, rollout stops if canary p95 latency exceeds 800 ms, a displayed count differs from the authorized dataset count, or an unauthorized tool call reaches execution. An access violation also needs incident response because rollback cannot undo an action already taken. The 5 percent share and one-week duration are example choices, not sufficient evidence by themselves. Too few requests or missing workload slices require more observation before full rollout.5
Production evidence feeds the next cycle. Reviewed failures become golden-set cases (Section 5.6), and user reports are checked against the requirements. A new model version, a changed requirement, or a missed SLO reopens the release decision record.
Book conclusion
A sufficient architecture meets the task’s requirements under tested conditions and records how it produced the result. Release review checks the complete outcome and the contributing components. Added complexity needs a required capability or a measured limitation, and live monitoring checks whether the system still meets the requirements after release.
Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1. MANAGE 4.1 covers post-deployment monitoring plans, including user input, appeal and override, decommissioning, incident response, recovery, and change management. The framework does not prescribe a monitoring method.↩︎
Warner, A., & Davidovič, Š. (2018). Canarying releases. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (Chapter 16). O’Reilly Media. https://sre.google/workbook/canarying-releases/. The chapter defines canarying as a partial, time-limited deployment of a change and its evaluation. Its examples come from Google services.↩︎
Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. https://doi.org/10.1017/9781108653985. Chapters 1 and 7 explain randomized experiments, their assumptions, metrics, and an overall evaluation criterion. The book’s examples come from large web services.↩︎
Jones, C., Wilkes, J., Murphy, N., & Smith, C. (2016). Service level objectives. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (Chapter 4). O’Reilly Media. https://sre.google/sre-book/service-level-objectives/. The chapter defines service level indicators and objectives and recommends latency percentiles over averages.↩︎
Warner, A., & Davidovič, Š. (2018). Canarying releases. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (Chapter 16). O’Reilly Media. https://sre.google/workbook/canarying-releases/. The chapter defines canarying as a partial, time-limited deployment of a change and its evaluation. Its examples come from Google services.↩︎