13  Evidence for release decisions

A candidate can give the right answer in a test and still use unauthorized data, repeat a tool call, or miss its latency target. A release review checks the complete result and the records of how the system produced it. The decision applies to one system version, its permitted inputs and actions, and the conditions tested.

Four section cards match Sections 13.1 to 13.4, including the four release-record subtopics. Optional components surround a model without an upgrade ladder. Records and controlled tests precede changes. An illustrative record checks count twenty-seven against twenty-seven, routing ninety-three of one hundred against an assumed ninety-five-of-one-hundred requirement, and twenty denied attempts, leading to HOLD. A separate tested-version PASS panel sits beside a limited rollout showing current and canary versions concurrently, separate SLO monitoring, and an orange stop or rollback route from canary to current.
Figure 13.1: Component choices follow task needs, and release decisions follow recorded checks. The illustrative 93-of-100 result fails the assumed 95-of-100 gate and stays on HOLD. A version that passes offline checks still needs separate live monitoring and a rollback route.

The map connects architecture choices, failure diagnosis, a release record, and monitoring after release. Each stage uses evidence from the model, retrieval, tools, and memory only where the selected system needs them.

13.1 Smallest sufficient architecture

The examples in Chapter 12 combine different components for different tasks. Adding all of them to a new product would add costs and failure paths without necessarily improving the result. The choices below start from the task and its tests, then add capabilities for requirements the current design cannot meet. This is the book’s engineering selection rule, not a fixed sequence that every product must follow.

  1. The task specification states the user-visible behavior, allowed inputs, required final state, non-goals, safety constraints, latency target, and cost target.
  2. Representative and adversarial examples, a baseline, and versioned evaluation records establish evidence before a design change.
  3. One model call is a baseline when the required information fits the request and no external action is needed. It is sufficient only if its measured results meet the task’s quality, safety, latency, and cost requirements.
  4. Deterministic context assembly supplies policy, history, examples, or structured records to the call without giving the model control over their selection.
  5. Retrieval supplies information from allowed external sources. Document search and structured queries solve different access problems, as Chapter 6 explains. Retrieval and generation remain separate evaluation targets.
  6. A workflow keeps known steps and branches in software, including branches selected from observations.
  7. An agent lets the model choose the next permitted tool action from observations when the useful path is not known in advance. Application code still limits actions and steps.
  8. Explicit planning makes dependencies and remaining work visible. MCP shares capability discovery and call formats. Multiple agents divide work among specialized loops. Long-term memory preserves selected information across sessions. Each is a separate choice justified by the task and its tests.

For example, a fixed workflow can query a database and validate the result without document retrieval or an agent. Adding MCP changes how an allowed tool is discovered and called, not whether the model needs to choose the next tool. The table relates each observed need to a likely design and the evidence needed to justify it.

Table 13.1: Architecture should grow in response to observed requirements.
Observed need Smallest likely pattern Evidence before escalation
Stable-rubric transformation of supplied text Prompted model call with structured output Golden set, schema validity, quality, latency, cost
Private or current document access RAG pipeline Naive baseline, retrieval labels, support and citation checks
Routing among known paths Deterministic or model-assisted workflow Route confusion matrix and branch tests
Iterative tool choice ReAct agent with a step limit Tool-selection tests, max-step fallback, repeated final-state success
Action dependencies need a revisable plan Plan-and-Act Plan validity, replan value, latency and cost delta
Independent expensive subtasks Dependency graph and parallel scheduler Dependency audit and critical-path measurement
Shared capabilities across hosts or teams MCP interface Measured integration burden, permission and transport tests
One agent hits tool, context, serial, or domain responsibility limits Orchestrator-worker or justified alternative Single-agent baseline and a metric for the measured limitation
Useful state must survive sessions Typed episodic or semantic memory Retrieval utility, correction, retention, and privacy tests

13.2 Failure localization before model changes

The selected architecture identifies the components that can affect a result. When an evaluation fails, replacing the model before identifying the cause can add unrelated output changes. A symptom narrows the investigation but does not prove which component failed. This table gives possible causes, records to inspect, and a controlled diagnostic test. If a fact was lost, the investigation follows it from the source through ingestion, retrieval, and the actual model request before judging generation.

Table 13.2: Symptoms guide trace inspection and controlled tests. They do not prove which component failed.
Symptom Possible causes First inspection Controlled next test
Answer ignores a source fact Ingestion, retrieval, context assembly, or generation Is the fact preserved in parsed text, indexed chunks, retrieval results, and the actual model request? Compare source, parsed text, index entry, retrieval result, and request snapshot before testing generation
Correct evidence, unsupported claim Generation or validation Does the claim follow from the supplied chunks? Faithfulness judge plus deterministic numeric check
Agent calls a plausible wrong tool Tool schema, router, model selection, or validation Are names, descriptions, arguments, and route classes distinct? Tool-choice golden set and description ablation
Agent loops Stopping policy or tool feedback Is progress represented in state and is the fallback reachable? Forced no-progress trace with step budget
Follow-up loses reference Workflow checkpoint or context assembly Was session state restored under the same ID? Restart and pronoun-resolution test
Old preference overrides correction Semantic-memory conflict policy Were both facts appended instead of superseded? Correction, precedence, and deletion test
Multi-agent output is incomplete Handoff, worker output, or join logic Did every required worker return a valid output, and did the join retain it? Missing-worker, malformed-output, partial-failure, and join-omission tests
MCP tool writes too broadly Authorization and server scope Which approved user intent, token permissions, and server access policy allowed the write? Denied-path, expired-token, and approval test

One-variable discipline

A controlled diagnostic changes one suspected component and states the expected metric change. It keeps the test cases and other relevant settings fixed, records the changed version, and checks effects on quality, safety, latency, and cost. A model change is justified when that investigation points to model behavior rather than a missing input or a failed control.

13.3 The release decision record

Section 13.2 established how to localize a failed test case. After an evaluation run, an accountable reviewer must decide whether the measured results and enforced controls support release.

  • Release decision record: A versioned document that links the task definition, evaluation results, component traces, thresholds, blocking failures, and accountable follow-up to a PASS or HOLD decision for specified operating conditions.

Reviewers use the release decision record to compare tested behavior with release requirements. A high aggregate score can conceal local failures: a model can produce a correct answer from unauthorized documents, or an agent can succeed inconsistently across repeated trials. Success on the last allowed attempt is still success if it meets the outcome, safety, latency, and cost requirements. The record places task, retrieval, control, and infrastructure evidence beside their thresholds. A change that alters tested behavior or risk requires a new record. Examples include adding a write-capable tool, broadening permissions, connecting a new MCP server, adding an agent route, storing a new memory category, or changing a safety threshold.

The record uses the evaluation suite introduced in Chapter 4 and execution traces like those in Chapter 12. If the golden set omits edge cases, reviewers may overestimate readiness. The following illustrative record extends the support-data example from Chapter 12 with release requirements and test results. Those results are assumed for the example, not measurements from a provider run or controls implemented by the shortened Chapter 12 listing.

A compact record contains system_version, task_scope, thresholds and slice results, source or index versions, tool permissions and budgets, protocol and memory versions when used, p95 latency, cost per case, the PASS or HOLD decision, blocking failures, the person accountable for follow-up, and the required work. The four sections below identify the evidence needed for those fields in the illustrative refund-count release.

Example: A release record leads to HOLD

The following illustrative record uses the refund-count trace from Sections 9.8 and 12.5. The tool count and displayed count are both 27, so that check passes. Intent-routing edge cases pass in 93 of 100 cases, below the required 95 of 100. Unauthorized tool attempts are rejected in 20 of 20 cases. The p95 latency is 760 ms against an 800 ms limit. Cost is $0.018 against a $0.020 limit.

Decision: HOLD. The routing slice fails a required threshold despite the passing count, control, latency, and cost checks. Repair the routing errors using development cases, then rerun the versioned validation suite. If the repair used failed validation cases, a fresh held-out sample is needed to check generalization. Twenty rejected unauthorized attempts show the result for those tests, not a guarantee for all possible calls. MCP, multiple agents, and persistent user profiles are not used in this example, so the record marks them not applicable.

13.3.1 Task and evaluation evidence

The release review begins with the evaluation lineage established in Chapter 4. A raw accuracy score does not explain what the system is allowed to do or whether it handles edge cases safely. The first section of the release record therefore defines success and summarizes the metric outcomes.

The task specification should list the final state, non-goals, and automatic failures. For the refund-count agent, the required final state is a count of permitted rows that matches the recorded dataset version and intent filter. Policy questions and technical-support requests are outside this route and need another permitted route or an out-of-scope response. The record includes baseline metrics, candidate results, and any judge calibration used for answers written in natural language. It breaks performance down by critical slices. One field fragment is task_scope: refund_count, intent_routing: 93/100, and required_intent_routing: 95/100. Separate results for normal, adversarial, and empty-input slices show which cases caused the shortfall. The sample size and threshold belong to this example’s release policy, not a universal readiness standard.

13.3.2 Retrieval and grounding evidence

Task evaluation checks the final answer against the requirements. The record should also show whether the system used authorized information of an appropriate version. The evidence path depends on the route: document retrieval follows indexed chunks through the model request to the answer and its citations, while a structured-data route follows access checks, filters, and tool results. Appendix A lists trace fields and operational evidence records.

An answer can look correct even when its trace omits source identity or freshness, or when application code read data outside the caller’s permissions. The release record traces each route’s information path. For the refund-count route, an illustrative field fragment records dataset_version: support-2026-09-01, caller_scope: approved_support_dataset, intent_filter: get_refund, row_ref: result-4821, and count: 27. Application code checks the caller’s permitted dataset before applying the intent filter. The value get_refund selects a category, not a permission. No document retrieval or citation is claimed. A policy-question route needs separate tests for ingestion and index freshness, the chunks preserved in the model request, supported citations, and access control.

13.3.3 Agent control evidence

The evidence path checks the information used. The control record must also show that application code permits only authorized actions and stops the loop within its configured budget. It documents state schemas, registered tools, permission checks, and execution limits.

Dynamic tool choice introduces execution risks that answer-only evaluation cannot expose. The release record lists every tool registered for the agent loop, including its schema, access checks, and error behavior. For the illustrative refund-count release, the field fragment is tools: [filter_by_intent, count_rows], dataset_write_tools: [], step_budget: 3, and unauthorized_attempts_rejected: 20/20. The filter tool still stores a temporary result record, as Section 9.2 explains. That record needs caller access checks and a retention rule even though the support dataset is unchanged. One step must have a defined meaning. In Section 12.5 it is a routing or tool attempt, so the successful path spends one step on routing and two on tools. A full controller also needs failure tests showing that invalid proposals consume the budget, that no further call executes at exhaustion, and that the caller receives a failure or escalation status. Unauthorized-caller and stale-version tests are separate from that stopping test. This is evidence required from the implementation, not behavior supplied by the release record itself.

13.3.4 Infrastructure evidence

Agent controls limit an individual loop. Shared servers, multiple agents, and persistent memory add controls that one loop’s tests cannot establish. The infrastructure section records the interfaces, access policies, and storage rules for the components actually used.

A controlled agent can operate without MCP, multiple agents, or long-term memory. The refund-count example records mcp: not_applicable, persistent_memory: not_applicable, p95_latency_ms: 760, and cost_per_case: 0.018. Latency needs a workload and measurement window, and cost needs a currency and a rule for counting model and tool charges. Here the cost is US dollars per evaluated case. If the system later uses MCP to access a billing server, the record identifies the protocol and server versions, the credential identity and permissions, and the allowed capabilities. It must not copy secret tokens into the record. Chapter 10 explains the version-dependent protocol details and the access checks that remain the host’s and server’s responsibility. If the system stores user preferences, deletion tests cover the primary store, derived indexes, caches, and any other retained copies required by its policy. Retrieved memory is data, not authority to change policy or execute a tool. Appendix D poses practice cases for diagnosis and release decisions, and Appendix E answers them. The References section lists the cited sources.

13.4 Operating rule for release decisions

An added component should meet a task requirement or address a measured limitation, with enough benefit to justify its costs and failure paths. Mandatory safety and access controls remain release conditions, not costs that an accuracy gain can outweigh. The decision record keeps task results beside information sources, tool and state traces, safety checks, latency, and cost. If a critical slice or required control fails, the result is HOLD even when the average score improves. A PASS supports only the tested version and operating conditions, with the uncertainty of the tests used. Later changes require a new comparison.

Passing offline tests does not establish a benefit on live traffic. If the candidate improves an offline metric, the product outcome identified in Section 4.1 is still a hypothesis to check after release. The NIST AI RMF lists post-deployment monitoring, user input, appeal and override, incident response, and change management among its outcomes (MANAGE 4.1).1 The following methods support different parts of that check:

  • Shadow deployment: The candidate processes copies of live requests while users receive only the current system’s answers. Comparing both outputs tests the candidate on real traffic. Tool calls with side effects are blocked or simulated.

  • Canary release: A partial and time-limited deployment of a change, evaluated before full rollout.2 Its stop rules are written before it starts.

  • Online controlled experiment: Random assignment of users or sessions to the current and candidate versions, followed by a comparison of product outcome metrics.3 It estimates the change in the chosen outcome under the experiment’s assignment and measurement conditions. Sample size, uncertain estimates, and interactions between groups can limit the conclusion.

  • Service level objective (SLO): A target value or range for a measured service indicator, such as p95 latency or error rate.4 The service policy defines the measurement window and the response to a breach, such as pausing rollout or rolling back.

Shadow deployment compares responses without exposing users to candidate decisions. A canary limits live exposure while monitoring failures. A controlled experiment estimates a product-outcome difference between assigned groups. These approaches can be combined, but none replaces access checks or the offline release requirements. A canary may also use randomized assignment. Monitoring reliability is a different question from measuring a product benefit.

After PASS, the refund-count agent could serve a selected 5 percent of analysts for one week while the remaining analysts use the current version. Before rollout, the record defines how analysts are assigned, the observation window and minimum request count for p95 latency, and who can stop the rollout. The platform measures the canary and current version separately. In this example, rollout stops if canary p95 latency exceeds 800 ms, a displayed count differs from the authorized dataset count, or an unauthorized tool call reaches execution. An access violation also needs incident response because rollback cannot undo an action already taken. The 5 percent share and one-week duration are example choices, not sufficient evidence by themselves. Too few requests or missing workload slices require more observation before full rollout.5

Production evidence feeds the next cycle. Reviewed failures become golden-set cases (Section 5.6), and user reports are checked against the requirements. A new model version, a changed requirement, or a missed SLO reopens the release decision record.

Book conclusion

A sufficient architecture meets the task’s requirements under tested conditions and records how it produced the result. Release review checks the complete outcome and the contributing components. Added complexity needs a required capability or a measured limitation, and live monitoring checks whether the system still meets the requirements after release.


  1. Tabassi, E. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1. MANAGE 4.1 covers post-deployment monitoring plans, including user input, appeal and override, decommissioning, incident response, recovery, and change management. The framework does not prescribe a monitoring method.↩︎

  2. Warner, A., & Davidovič, Š. (2018). Canarying releases. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (Chapter 16). O’Reilly Media. https://sre.google/workbook/canarying-releases/. The chapter defines canarying as a partial, time-limited deployment of a change and its evaluation. Its examples come from Google services.↩︎

  3. Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. https://doi.org/10.1017/9781108653985. Chapters 1 and 7 explain randomized experiments, their assumptions, metrics, and an overall evaluation criterion. The book’s examples come from large web services.↩︎

  4. Jones, C., Wilkes, J., Murphy, N., & Smith, C. (2016). Service level objectives. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (Chapter 4). O’Reilly Media. https://sre.google/sre-book/service-level-objectives/. The chapter defines service level indicators and objectives and recommends latency percentiles over averages.↩︎

  5. Warner, A., & Davidovič, Š. (2018). Canarying releases. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (Chapter 16). O’Reilly Media. https://sre.google/workbook/canarying-releases/. The chapter defines canarying as a partial, time-limited deployment of a change and its evaluation. Its examples come from Google services.↩︎