8  Operating a grounded-answer system end to end

Suppose that after a routine weekly update, answers about newly added policies start citing the wrong documents or none at all, while latency and error rates stay normal. The groundedness defined in the Introduction has fallen even though the service is available. The preceding chapters produced versioned artifacts, recoverable distributed jobs, durable storage, serving bundles, a versioned retriever, evaluation evidence, and retryable pipelines. This chapter operates them as one system whose release evidence connects such an alert to a verified recovery.

The exact release, retrieved passages, and build run must be connected before a failed answer can lead to a trustworthy repair. The map follows those links through diagnosis, release checks, and the next improvement cycle.

Chapter map: operating a grounded-answer system end to end

The sections answer four linked questions:

  • 8.1–8.3: What must the service guarantee, how is the latency budget divided, and how do documents become a validated retrieval generation?
  • 8.4–8.6: How are model runs compared, how are a serving bundle and a retrieval collection released together, and how do request traces connect to population quality?
  • 8.7–8.8: How is a groundedness incident traced to its cause, and where does the control belong that stops the same failure next time?
  • 8.9–8.10: Which evidence decides a release, and how does production feedback start the next improvement cycle?
Ten panels connect acceptance targets, versioned artifacts, committed retrieval generation, reproducible experiments and compatible releases to trace-linked quality, evidence-led incident diagnosis, regression gates, release decisions and a measured canary feedback cycle.
Figure 8.1: Versioned artifacts and release evidence connect a complete build to one compatible release, an evidence-led groundedness diagnosis, prevention gates, and the next measured improvement cycle.

8.1 Defining the end-to-end service target

Every technical layer is now available. An end-to-end design still needs a product definition that prevents teams from optimizing local metrics while breaking trust, privacy, or reproducibility. The workload, evidence responsibilities, and system guarantees define how the components fit together.

The proposed case is a multi-tenant grounded-answer service. Its guarantees below are acceptance targets, not properties supplied automatically by retrieval and generation. Users ask questions about a governed document corpus. The system retrieves authorized passages, generates an answer, attaches citations, and records enough evidence to evaluate and replay the request without storing unnecessary sensitive content.1

The live path filters permissions, retrieves passages and generates a cited candidate answer with a trace. A separate sampled offline path checks claims and citations and feeds quality evidence into a release gate.
Figure 8.2: Authorized retrieval supplies passages to a candidate answer, not proof of groundedness. A separate sampled evaluation checks claims and citations and contributes evidence to release decisions; live requests do not all wait for that judge.

The retriever selects and authorizes evidence. The generator constructs an answer from that context, while the evaluator checks claims and citations. The release identity records the exact combination. Together, these records let investigators localize a failure without labeling every wrong answer as a model-weight defect.

Table 8.1: Acceptance targets for the proposed service. Passing evidence is required; these are not automatic RAG properties.
Guarantee Meaning Evidence
Authorization precedes use A user may retrieve and cite only documents permitted at query time. Tenant/ACL filter version and retrieved source IDs
Claims are grounded Material factual claims are supported by supplied, cited context. Claim-level judge output and citation verifier
Release configuration is reconstructable Recorded and retained compatible identities support reconstruction, not universal numerical equality. Serving-bundle and collection manifests
Derived state is rebuildable Retained source versions, model, transforms, and index settings support a tested rebuild. Chunk, embedding, and index build manifests
Promotion is reversible A retained compatible release can be restored only if current access and deletion policy still permits it. Rollback drill and retained versions
Telemetry is governed Tracing preserves diagnostic identifiers while applying retention and redaction policy. Sampling/redaction configuration and audit

Each system guarantee has a producing component, a machine-verifiable record, and a stop action. For example, the retrieval pipeline produces ACL-aware source identities, the trace records what was selected, and the release gate blocks a candidate whose negative-access tests fail. A guarantee without a stop action is only an aspiration.

The evaluator checks citation support independently of answer correctness.2 Current permissions and deletion policy apply at every retrieval, source-fetch, citation-preview, cache, and rollback path, even if an older corpus snapshot allowed access.3 Full replay is possible only within the permitted retention window and with compatible retained inputs.

The latency objective also has to be divided among components. The grounded-answer path runs several steps before the model produces its first token, so retrieval, reranking, and prefill share one TTFT budget. A deterministic planning allocation is an intentionally chosen time allowance for one component in that design budget. It is not a percentile of the end-to-end latency distribution. Per-stage p95 observations can inform those allocations, but their sum does not determine end-to-end p95. The team compares end-to-end TTFT measurements under the expected workload with the p95 objective, or validates a statistical model of the joint stage latencies, before deciding whether the objective is met.

Example: time to first token for one grounded answer

The service objective is p95 TTFT \(\leq\) 1,200 ms. Measured per-stage p95 values inform the following chosen planning allocations:

  • Before retrieval: Gateway and authorization 20 ms, query embedding 15 ms.
  • Retrieval (Chapter 5): Filtered dense and sparse search 60 ms, reranking 40 candidates 120 ms, packing 5 ms.
  • Engine (Chapter 4): Queue 150 ms, prefill of a 6,500-token prompt 600 ms.
  • Planned total: The chosen allocations sum to 20 + 15 + 60 + 120 + 5 + 150 + 600 = 970 ms, leaving 230 ms of allocation headroom.
  • Option A: Reranking 80 candidates instead of 40 adds about 120 ms, for a 1,090 ms allocation. It fits this allocation, with 110 ms of allocation headroom.
  • Option B: An answer-groundedness judge requires the generated answer. Holding the stream for that judgment adds answer-generation time plus about 900 ms for the judge. Even if answer-generation time is omitted, the planning allocation is already about 1,870 ms, above the objective. The judge therefore runs on sampled traffic after the answer is streamed.
  • Conclusion: Deeper retrieval and inline evaluation compete for the same design allocation. The team compares end-to-end TTFT measurements under the expected workload with the p95 objective to decide whether either request path meets it.

Release observation window: collecting enough evidence before promotion

A candidate is promoted only after its evidence covers the traffic patterns that matter:

  • Baseline: Compare a candidate with the last known-good release on the same request mix and protected quality slices.
  • Window: Collect enough shadow or canary traffic to include peak and off-peak behavior, cache warm-up, retries, delayed evaluator labels, and one planned rollback rehearsal.
  • Measures: Track availability, TTFT and ITL, retrieval recall and freshness, groundedness, authorization failures, cost, and evaluator disagreement with release identity attached.
  • Decision: Promote only when every required slice and guardrail has data for the full window. Otherwise, extend observation, repair the candidate, or keep the known-good release.4
  • Conclusion: A defined measurement window prevents the team from accepting a short favorable sample as complete release evidence.

8.2 Coordinating artifact versions and release responsibilities

The product guarantees are explicit. Incidents become untraceable when ‘the model,’ ‘the data,’ or ‘the index’ refers to a mutable alias instead of a specific generation. The release records which run built each artifact and which later request or deployment consumed it.5

Table 8.2: Immutable identities form the lineage path from a request back to source evidence and build runs.
Artifact Example immutable identity Owning layer Required references
Source corpus corpus_manifest=sha256:... Data governance/storage Source object versions and ACL snapshot
Chunk set chunker:v4/corpus:41 Retrieval pipeline Corpus, code, parameters
Embedding set embed:e3/chunks:41 Retrieval pipeline Chunks, model revision, normalization
ANN collection collection:c41-e3-i7 Vector platform Embedding set and index config
Training run experiment:ft-2026-071 Training/experiment tracking Dataset, base model, code, config, image
Model checkpoint model:answerer-v12 Model registry Training run and checkpoint manifest
Serving bundle bundle:b12.4 Inference platform Model, tokenizer, engine, quantization, prompt defaults
Evaluation report eval:grounded-v9/b12.4/c41 Evaluation platform Suite, judge, bundle, collection
Deployment release:r218 Control plane Bundle, resolved collection and query encoder, report, approval
Request trace trace_id + release:r218 Observability Runtime spans, retrieved IDs, output digest

The release descriptor introduced in Section 5.5 now identifies the serving bundle, retrieval collection, query encoder, prompt and policy configuration, and evaluation evidence used together for one request or run. A candidate can have a descriptor before approval. The descriptor identifies the combination, while the release decision determines whether it may receive traffic.

Artifact governance: immutable resolution of routing aliases

Separating mutable routing aliases from immutable artifacts preserves experiment reproducibility:

  • Human-friendly aliases such as production, current-corpus, or latest-model are mutable routing pointers and therefore unsuitable as unresolved experiment inputs.6
  • At entry, a run or request resolves one immutable release descriptor and carries that same bundle, collection, query encoder, and policy-compatible configuration across retries. In-flight requests continue using their resolved combination unless current policy requires them to stop.

For one failed request, the trace identifies the deployment. The deployment identifies the bundle, collection, and approved evaluation, which in turn identify their build runs and immutable inputs. In the release direction, a candidate receives a deployment ID only after every required upstream identity is present.

Release governance: responsibilities and emergency escalation

Defining approval roles and escalation authorities prevents unvetted production changes:

  • Producing teams are responsible for the completeness and evidence of the artifacts they produce. The release owner assembles the compatibility record, and the operations owner manages rollout, rollback, and the evidence window.
  • Approval records the candidate, evidence links, accountable person, and expiry or review point. Escalation identifies who can stop traffic when a hard guardrail fails. It is not replaced by a favorable average score.

8.3 Tracing documents from ingestion to retrieval

Those recorded links define the required connections. The first active path turns governed documents into searchable state without losing source, access, deletion, or version information. Committed generations and quality checks make the ingestion path reproducible.

The document ingestion path progresses through eight verifiable stages.

  1. Capture the source. A source snapshot records the document versions, access metadata, and deletion watermark in an immutable corpus manifest. This is the input that later checks compare against.
  2. Validate the intake. Validation checks file types, character decoding, duplicate rate, policy compliance, and item counts. An invalid batch stops here rather than producing searchable content.
  3. Create traceable chunks. A versioned splitter records each chunk’s content hash, access metadata, and a locator that can return a reader to the source. PDF pages, OCR regions, and text spans can need different locators. Byte offsets are used only when an extraction and normalization mapping has been tested.
  4. Produce compatible vectors. The embedding task records the model revision and normalization rule with the vectors. A later query must use the compatible query representation.
  5. Build a candidate index. The index task creates an approximate-nearest-neighbor (ANN) structure in a staging collection. Live traffic continues to use the current production alias.
  6. Reconcile the candidate. A coordinator compares vector counts, unique IDs, deletions, and partition coverage with the corpus manifest. A missing partition or an unprocessed deletion leaves the candidate incomplete.
  7. Evaluate the result. Benchmark retrieval cases run under expected load to check recall, latency, and access-filter behavior.
  8. Publish one committed generation. Only the coordinator records a durable committed-generation manifest and makes the candidate eligible for canary traffic. A successful worker process alone does not make the collection visible.

The following checks define completion for the candidate generation instead of relying on process exit status.7

Table 8.3: Completion checks. Matching counts and hashes do not by themselves establish parsing correctness, semantic coverage, or authorization.
Gate Pass criterion Why exit code is insufficient
Completeness Expected source/chunk/vector counts reconcile by tenant and partition A worker can skip a partition and still exit successfully
Integrity Hashes, dimensions, metadata schema, and unique IDs match the schema definition Serialization may succeed with wrong shape or duplicates
Authorization ACL mutation and negative-access tests pass Readable vectors can still expose forbidden content
Retrieval quality Required recall and citation coverage pass on held-out slices An index can be healthy but semantically poor
Freshness Inserted and updated content becomes searchable, and deleted content stops being returned, within the objective Static benchmark ignores lifecycle lag
Recovery Candidate restores or rebuilds within RTO Replication status does not prove usable recovery

The completion manifest is the commit record. Expensive workers may retry and write attempt-scoped shards, but only the coordinator can convert those outputs into a consumable generation after reconciliation. The promotion task reads that committed evidence instead of inferring completeness from worker status alone.

The candidate must remain immutable between checking and use. The conditional publication rules from Sections 3.4 and 7.6 reject old coordinators and delayed attempts. Collection-alias atomicity alone does not make a serving, retrieval, prompt, and cache change one transaction.8

8.4 Reproducible model experiments

The retrieval collection can now be reproduced. The generator or judge model may still need fine-tuning on examples from the company’s own documents, for example to follow the citation format reliably. Such a run can train a LoRA adapter (Section 4.6) or all weights, and when it is long and distributed it must survive interruption without losing experiment identity. One run manifest connects distributed strategy, scheduling, checkpointing, and experiment evidence.

Model-state and activation memory decide whether DDP, FSDP, tensor, pipeline, sequence/context, or expert parallelism, with or without activation recomputation, fits the constraint (Section 2.4). A LoRA run holds optimizer state only for the adapter, so it often fits where full fine-tuning would need sharding. The communication-heavy dimension uses the fastest available fabric, and the scheduler admits the complete topology. Each rank records world, node, local-rank, device, and shard identity.

Table 8.4: A distributed model run is complete only when its artifacts and evidence are recoverable and traceable.
Stage Required identity/evidence Failure point
Allocation Code commit, image digest, data manifest, base model, config, topology Rendezvous and gang allocation
Active training Loss, throughput, MFU, memory, power, communication, data time Rank failure, straggler, NaN, device error
Checkpoint All model/optimizer/RNG/sharding state and completion manifest Partial save or lost node-local state
Resume Checkpoint generation plus compatible code/config/topology rules Silent restart from wrong step or sampler position
Held-out evaluation Held-out suite, raw outputs, evaluator identity, quality/cost slices Metric-only success without failed cases
Register Validated immutable model version and lineage Mutable path or incomplete bundle

The distributed checkpoint saves the state supplied by the training application, subject to format and version compatibility.9 The experiment tracker compares candidate runs, while the orchestrator records task execution and retry history. The checkpoint manifest links both. Model registration occurs in a separate gate after held-out evaluation.

Reproducibility audit: identifying drift between twin execution runs

Comparing execution manifests reveals configuration drift and non-deterministic environmental factors:

  • Same baseline: Both runs use the same dataset snapshot, evaluation suite, tokenizer, and serving test mix.
  • Changed factors: Run B changes the sharding strategy and image digest. The manifest records those differences explicitly.
  • Comparison: Run B improves throughput but increases recovery time and shows a quality drop on the long-context slice.
  • Decision: The candidate is not promoted until the long-context regression is explained or repaired. The aggregate throughput gain alone is insufficient.
  • Conclusion: Comparability comes from holding the declared inputs constant and making every changed factor visible.

8.5 Releasing serving and retrieval changes together

Model and collection artifacts are independently versioned. A grounded-answer response depends on their compatibility with the prompt, citation formatter, tokenizer, and evaluator, so promoting only one alias can create an untested combination. The deployable unit is one immutable release descriptor whose component combination has passed the linked artifact-pairing checks in Table 8.2, the staged evidence in Table 8.5, and the compatible encoder and collection relation in Section 5.5.

  • Release candidate: One immutable serving bundle, retrieval collection, prompt/policy configuration, and evaluation report proposed for production traffic.

Each test stage adds realism and risk after the candidate has passed the preceding evidence check.

Table 8.5: Release evidence progresses from deterministic compatibility to production-like and controlled live behavior.
Test stage Traffic Checks Promotion condition
Interface Synthetic Startup, API schema, tokenizer, tool and citation parsing All deterministic tests pass
Fixed benchmark Recorded prompts and lengths TTFT, ITL, throughput, memory, output format No required regression
Offline quality Held-out cases Groundedness, citations, safety, retrieval quality Slice thresholds and uncertainty pass
Shadow Copied live inputs, no user-visible result Load mix, routing, trace completeness, disagreement No unexplained severe divergence
Canary Small controlled live share SLOs, errors, evaluator proxies, complaints, cost Evidence window and guardrails pass
Production Ramped share Same guardrails plus capacity and rollback readiness Stable after each ramp step

The gateway stamps the resolved release, bundle, collection, prompt, and policy identities into request context. The selected router policy uses capacity and reusable KV prefixes, subject to model/configuration compatibility and the allowed tenant boundary.10 The engine records queue, prefill, decode, token, memory, and error evidence. Retrieval spans record filters, candidate stages, selected source IDs, and score conventions.

8.6 Linking request traces to population quality

The release can receive production traffic. Operators need both aggregate alerts and one replayable request path, while evaluators need enough governed context to judge behavior. Service telemetry, AI traces, sampled judgments, delayed labels, and release identities form one evidence path.11

Table 8.6: Proposed aggregate/request evidence joins. Correlation supports investigation, and replay depends on permitted retained payloads.
Evidence layer Population signal Request-level evidence
Gateway/router Rate, errors, queue, retries, replica balance Tenant class, release ID, route and retry spans
Retriever Recall proxy, no-result rate, score and freshness drift Query digest, filters, candidate IDs/scores, selected chunks
LLM engine TTFT, ITL, throughput, KV pressure, batch size Prompt/output token counts, replica, cache reuse, timing
Quality Groundedness/citation/safety rates with uncertainty Raw evaluator result, claim labels, parse status
Feedback Complaint and acceptance rates by slice Linked feedback type and timestamp
Cost Tokens, GPU time, retrieval/rerank calls, judge sampling Per-request attributed usage

The telemetry owner applies the data policy when selecting trace samples and deciding which request fields remain available for replay. Aggregate metrics can remain longer than trace bodies, while trace bodies need a defined retention period and access rule.

Opaque IDs and hashes can still reveal sensitive information, especially when the original value is predictable. Redaction, encryption, access, and retention are implemented controls, not guarantees supplied by identifiers or telemetry conventions. When a payload expires, its identity may support lineage inspection but not full replay.12 Hashed or opaque source identities replace full content where possible, secrets and personal data are redacted before export, governed payloads are encrypted and access-controlled, and trace bodies expire independently of low-cardinality metrics. The evaluator revision and raw parse outcome help distinguish a judge change from a product change. The record also binds the prompt, rubric, model settings, preprocessing, and reference set. Controlled rescoring and human review of critical disagreements remain necessary.13

8.7 Finding the cause of a groundedness incident

The evidence path spans every component. The incident begins with an available service whose answers for newly added policy documents become less grounded. An ordered investigation localizes the failing artifact or stage before any component is blamed.

The following incident is a constructed example. Its numbers illustrate the diagnosis and recovery sequence, not results from an external production system.

Regression analysis: quantifying groundedness loss across document slices

Slicing benchmark evaluations pinpoints quality regressions introduced by updated retrieval collections:

  • Production change: Release r218 keeps serving bundle b12.4 but moves retrieval alias from collection c40-e3-i7 to c41-e3-i7.
  • Alert: The sampled groundedness rate falls from 0.91 to 0.78 for the new-policy slice. HTTP errors and TTFT remain within objective.
  • Trace evidence: Failed answers have low retrieval scores and omit source IDs introduced in corpus generation 41.
  • Lineage check: The corpus manifest expects 12.0 million chunks, but the committed collection contains 11.2 million. One tenant partition is absent.
  • Pipeline evidence: An embedding task retried after a worker loss. Its completion wrapper reported process exit zero as success, while the release coordinator did not reconcile per-partition counts before promotion.
  • Immediate mitigation: The release returns to the compatible c40-e3-i7 and b12.4 combination only after current deletion and access checks pass, affected owners receive notification, and failed traces remain preserved.
  • Repaired candidate: A new candidate contains the rebuilt partition, reconciled counts and hashes, authorization/retrieval/freshness results, and shadow and canary evidence.
  • Prevention: Completion-manifest reconciliation becomes a hard task gate, and compute tasks no longer change aliases.
  • Conclusion: Availability and latency stayed normal while answer quality fell. Linked release, trace, storage, and pipeline records identified the incomplete retrieval collection.
Table 8.7: Hypotheses are ranked by correlated evidence rather than component familiarity.
Hypothesis Evidence that weakens or supports it
LLM model regressed Weak: serving-bundle identity did not change, failures cluster by corpus slice.
Inference overloaded Weak: queue, TTFT, ITL, memory, and error distributions remain stable.
Embedding model incompatible Weak: both old and candidate collections use embedding version e3.
Access filter removed evidence Possible until negative/positive ACL traces show correct filter decisions.
Index is incomplete Strong: expected/actual count mismatch and missing partition IDs align with failed cases.
Judge changed Weak: evaluator version is constant and user complaints corroborate the same slice.

8.8 Regression prevention with evaluation gates

Rollback can restore the previous answer quality for requests covered by the known-good collection, provided that the older state still meets current access and deletion rules. Rebuilding one missing partition fixes this generation but leaves the pipeline capable of promoting another incomplete collection. A prevention control at the visibility point stops another invalid collection from reaching live traffic.

The prevention plan for this constructed incident uses seven connected controls. It does not claim that seven controls are necessary or sufficient for every system.

  1. Check the candidate before promotion. The verification coordinator checks partition row counts, unique IDs, vector dimensions, and access-control parity. This blocks an incomplete collection before it becomes visible.
  2. Keep writers away from the live alias. Compute workers write only to a candidate staging namespace. Only the verification coordinator has permission to change the live alias. Compute-worker identities cannot publish their partial output.
  3. Require a durable success record. The coordinator accepts a completion manifest tied to the reconciled candidate, not only a zero exit code. Hashes and signatures can add integrity checks when the threat model requires them, but they do not prove that the candidate is complete or authorized.
  4. Keep the failed case in the release suite. A targeted regression slice taken from the production incident gives the release gate a case that detects the same failure.
  5. Test recovery before an incident. A rehearsal removes a partition from a candidate and verifies that promotion stops without changing the live alias.
  6. Watch coverage after promotion. Searchable-count and partition-freshness signals reveal an unexpected loss of available content.
  7. Review canary evidence. Groundedness telemetry during the stabilization window checks whether the released system remains within its release objectives.

Release gate principle: optimal control placement

Placing preventive gates before deployment detects defects before user impact:

  • The preventive gate belongs at the point where invalid state could become visible.
  • A dashboard can detect a bad collection after promotion, while a commit manifest and promotion check block the invalid collection from live traffic.

8.9 Deciding release readiness from recorded evidence

The incident has produced a durable prevention control. Future releases need a compact cross-layer checklist so evidence is assembled before the rollout window rather than during an incident. Each responsible component contributes a verifiable artifact or measurement to the candidate release.

Table 8.8: A release combines quality, performance, safety, reproducibility, recovery, and responsibility evidence.
Layer Release evidence Stop condition
Artifacts/lineage Immutable model, tokenizer, image, prompt, data, collection, config, evaluator IDs Any unresolved mutable dependency
Training (when used) Run manifest, checkpoints, held-out quality, MFU/throughput, recovery test Missing state or unexplained instability
Serving Compatibility, load, soak, cold-start, capacity, rollback results SLO or rollback failure
Observability Correlated logs/metrics/traces, dashboards, alerts, redaction/retention Blind request stage or ungoverned payload
Evaluation Calibrated judge, raw results, consistency, slices, human adjudication Unstable or unvalidated critical slice
Storage/retrieval Completeness, integrity, ACL, freshness, recall, tail latency, restore Count/hash/authorization mismatch
Pipeline Idempotency, retry/backfill tests, task descriptions, promotion gate Side effect can repeat or partial output can commit
Operations Responsible person, runbook, canary plan, capacity, incident and rollback authority No accountable decision path

A stop condition has priority over a favorable aggregate score.

The release policy treats retrieved documents as untrusted evidence, not privileged instructions. Authorized content can still contain indirect prompt injection: instructions embedded in external content that try to redirect the application.14 A retrieved document can contain text that tries to override application instructions, disclose protected information, or influence a tool call. Authorization permits retrieval of the document. It does not authorize the model or application to obey text that conflicts with its governing instructions and tool rules. Adversarial retrieval and evaluator cases test refusal and containment where untrusted content enters the request. No single filtering or prompting rule is assumed to solve the attack class. High overall groundedness cannot compensate for an authorization failure, and high throughput cannot compensate for an untested rollback path. The release owner resolves each red condition explicitly: failed authorization, a missing required artifact, an unmet evaluation threshold, an incomplete rollback test, or an unresolved canary stop condition. Absence of evidence is not a pass.

Operations governance: approval policy versus escalation authority

Dividing policy approval from emergency escalation enables rapid incident response:

  • Approval confirms that the recorded evidence satisfies the release policy and names the accountable release owner.
  • Escalation covers a live stop: The on-call or incident authority can pause traffic, restore the known-good aliases, preserve evidence, and notify the affected owners while the release decision is revisited.

8.10 Improving the system from production feedback

The checklist can approve one candidate. Production data, traffic, providers, model aliases, document corpora, hardware, and costs continue to change after release. Each release is a measured hypothesis, and its failures become inputs to the next reproducible evaluation and pipeline run.

The operating loop is now complete. For example, a groundedness alert on newly added documents signals a change in measured groundedness for that document slice. A correlated trace supplies a replay case, the evaluation suite checks the repaired candidate, and the pipeline records the resulting serving and retrieval pair. Canary evidence then supports promotion or returns traffic to the known-good pair. The same evidence path supports that investigation. Fixed evaluator identity, controlled rescoring, and adjudication are needed to distinguish a product change from judge instability before the next release decision.

Development and incident-regression cases stay distinct from independently held-out evaluation. User acceptance is not automatically a correct label, and feedback can change future data in harmful ways. Improvement must be measured rather than assumed from completing the loop.15

Table 8.9: Improvement loop. Replay and measured improvement remain conditional, not guaranteed by recorded identities.
Question Evidence source Typical response
Is the service available? SLO metrics, errors, traces, capacity Serving-path operation and capacity adjustment
Is the behavior good? Evaluators, human labels, complaints, business outcomes Failure triage by slice
Why did it happen? Trace-to-bundle-to-run-to-artifact lineage Hypotheses for investigation and controlled replay
Can it be reproduced? Permitted retained inputs, images, configs, seeds, and raw outputs Replay within the retention and runtime compatibility boundary
Is the change better? Held-out and production-like comparative evidence Candidate approval, rejection, or revision
Can it be released safely? Compatibility, canary, rollback, recovery, policy gates Controlled promotion
Did the fix persist? Post-release guardrails and incident slice Closure after the evidence window

Chapter 8 summary

Key system integration and operational principles established in this chapter:

  • Core mechanisms: A release descriptor pins one tested serving bundle, retrieval collection, prompt, and policy. Request traces carry those identities, so a groundedness alert can be followed to a collection build and its pipeline run.
  • Governing trade-offs: Retrieval depth, reranking, and inline evaluation compete for one TTFT budget. A longer observation window gives stronger release evidence but slows delivery.
  • Failure modes & defenses: An incomplete collection passed because a worker’s exit code counted as success. Reconciling counts before promotion, keeping compute tasks away from the live alias, and adding the incident to the release suite stop the same failure before users see it.

Chapter checkpoint

Review Questions 39–41 in Appendix B, Section B.1, to test incident diagnosis, the placement of the prevention gate, and the end-to-end latency budget.

System synthesis: reliable AI platform engineering principles

Unified platform engineering principles connecting the complete machine learning lifecycle:

  • In this platform design, reliable operation depends on explicit connections among versioned artifacts, compute strategies, serving stages, telemetry, evaluation, durable state, and orchestration.
  • For every result, the operating model identifies its exact inputs, responsible component, validation evidence, failure behavior, reproduction path, and rollback path.

  1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. Curran Associates. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html. The original RAG paper combines retrieval with generation in tested QA/generation tasks. Authorization, citation checking, and replay are additional book requirements, not guarantees established by that paper.↩︎

  2. Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations. In Proceedings of EMNLP 2023 (pp. 6465–6488). Association for Computational Linguistics. DOI: 10.18653/v1/2023.emnlp-main.398. https://aclanthology.org/2023.emnlp-main.398/. ALCE measures correctness and citation quality separately and finds incomplete support in its tested systems. Source entailment does not establish that a source is true, current, or authorized.↩︎

  3. Microsoft. (2026, August 24). Security filters for trimming results in Azure AI Search (Examples use API 2026-04-01). https://learn.microsoft.com/en-us/azure/search/search-security-trimming-for-azure-search. The official security-filter pattern requires every query to use the relevant permissions and is not itself authentication. The surrounding system also applies current-policy checks to derived caches and rollback.↩︎

  4. Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data (pp. 1123–1132). IEEE. DOI: 10.1109/BigData.2017.8258038. https://storage.googleapis.com/gweb-research2023-media/pubtools/4156.pdf. The ML Test Score supports important-slice, pre-serving, canary, monitoring, and rollback checks. A full time window alone is not a sample-size or uncertainty guarantee.↩︎

  5. Moreau, L., & Missier, P. (Eds.). (2013, April 30). PROV-DM: The PROV data model. W3C Recommendation. https://www.w3.org/TR/2013/REC-prov-dm-20130430/. PROV-DM represents entity/activity generation, use, and derivation. It does not verify the truth of links or prescribe this release workflow.↩︎

  6. MLflow contributors. (n.d.). MLflow model registry (MLflow 2.4.2 archived documentation). https://www.mlflow.org/docs/2.4.2/model-registry.html. MLflow 2.4.2 explicitly permits alias reassignment. An alias can be accepted at entry if resolved once to a retained immutable version. Repeated resolution is the unsafe case.↩︎

  7. Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781–1794. DOI: 10.14778/3229863.3229867. https://www.vldb.org/pvldb/vol11/p1781-schelter.pdf. The original data-quality system supports declared constraints and aggregation-based checks. Matching totals and hashes do not establish semantic completeness, correct source locators, or ACL propagation.↩︎

  8. Qdrant. (n.d.). Collections (Collection aliases and switching). https://qdrant.tech/documentation/manage-data/collections/. Qdrant’s atomic grouped alias update is confined to its API. The broader request-pinning and coordinator protocol remains a book design.↩︎

  9. PyTorch Contributors. (2025, June 16). Distributed Checkpoint: torch.distributed.checkpoint. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/distributed.checkpoint.html. The distributed-checkpoint API saves supplied state with format/version constraints. The surrounding workflow adds explicit sampler, scheduler, scaler, step, manifest, and registration requirements.↩︎

  10. NVIDIA. (n.d.). Routing concepts (Dynamo 1.0.2 documentation). https://docs.nvidia.com/dynamo/v1.0.2/components/router/routing-concepts. Dynamo 1.0.2 gives one KV-aware routing mode with load/cache costs. Other routers and modes differ, and the citation is not an unconditional latency or isolation guarantee.↩︎

  11. OpenTelemetry authors. (2026, January 14). Traces (OpenTelemetry documentation). https://opentelemetry.io/docs/concepts/signals/traces/. OTel supplies span context, attributes, and links. Joining releases and delayed judgments requires application instrumentation and does not by itself prove causation or full replayability.↩︎

  12. OpenTelemetry authors. (2026, January 14). Handling sensitive data (OpenTelemetry security guidance). https://opentelemetry.io/docs/security/handling-sensitive-data/. OTel recommends minimization and redaction and warns that predictable inputs can be recovered from hashes. Retention, encryption, and authorization still need implementation and policy.↩︎

  13. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena (arXiv:2306.05685v4). arXiv. Author manuscript. The judge study documents bias and reasoning limits in preference evaluation. Stable identity alone does not make repeated judgments stable or establish a groundedness error rate.↩︎

  14. Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection (arXiv:2302.12173v2). arXiv. Author manuscript. The original study demonstrates attacker instructions in external content affecting tested LLM-integrated applications. It does not establish that a particular defense completely solves the problem.↩︎

  15. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503–2511. https://papers.nips.cc/paper_files/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf. Hidden Technical Debt Section 4 explains direct and hidden feedback that changes future data. A feedback loop can worsen behavior rather than automatically improve it.↩︎