8 Operating a grounded-answer system end to end
Suppose that after a routine weekly update, answers about newly added policies start citing the wrong documents or none at all, while latency and error rates stay normal. The groundedness defined in the Introduction has fallen even though the service is available. The preceding chapters produced versioned artifacts, recoverable distributed jobs, durable storage, serving bundles, a versioned retriever, evaluation evidence, and retryable pipelines. This chapter operates them as one system whose release evidence connects such an alert to a verified recovery.
The exact release, retrieved passages, and build run must be connected before a failed answer can lead to a trustworthy repair. The map follows those links through diagnosis, release checks, and the next improvement cycle.
Chapter map: operating a grounded-answer system end to end
The sections answer four linked questions:
- 8.1–8.3: What must the service guarantee, how is the latency budget divided, and how do documents become a validated retrieval generation?
- 8.4–8.6: How are model runs compared, how are a serving bundle and a retrieval collection released together, and how do request traces connect to population quality?
- 8.7–8.8: How is a groundedness incident traced to its cause, and where does the control belong that stops the same failure next time?
- 8.9–8.10: Which evidence decides a release, and how does production feedback start the next improvement cycle?
8.1 Defining the end-to-end service target
Every technical layer is now available. An end-to-end design still needs a product definition that prevents teams from optimizing local metrics while breaking trust, privacy, or reproducibility. The workload, evidence responsibilities, and system guarantees define how the components fit together.
The proposed case is a multi-tenant grounded-answer service. Its guarantees below are acceptance targets, not properties supplied automatically by retrieval and generation. Users ask questions about a governed document corpus. The system retrieves authorized passages, generates an answer, attaches citations, and records enough evidence to evaluate and replay the request without storing unnecessary sensitive content.1
The retriever selects and authorizes evidence. The generator constructs an answer from that context, while the evaluator checks claims and citations. The release identity records the exact combination. Together, these records let investigators localize a failure without labeling every wrong answer as a model-weight defect.
| Guarantee | Meaning | Evidence |
|---|---|---|
| Authorization precedes use | A user may retrieve and cite only documents permitted at query time. | Tenant/ACL filter version and retrieved source IDs |
| Claims are grounded | Material factual claims are supported by supplied, cited context. | Claim-level judge output and citation verifier |
| Release configuration is reconstructable | Recorded and retained compatible identities support reconstruction, not universal numerical equality. | Serving-bundle and collection manifests |
| Derived state is rebuildable | Retained source versions, model, transforms, and index settings support a tested rebuild. | Chunk, embedding, and index build manifests |
| Promotion is reversible | A retained compatible release can be restored only if current access and deletion policy still permits it. | Rollback drill and retained versions |
| Telemetry is governed | Tracing preserves diagnostic identifiers while applying retention and redaction policy. | Sampling/redaction configuration and audit |
Each system guarantee has a producing component, a machine-verifiable record, and a stop action. For example, the retrieval pipeline produces ACL-aware source identities, the trace records what was selected, and the release gate blocks a candidate whose negative-access tests fail. A guarantee without a stop action is only an aspiration.
The evaluator checks citation support independently of answer correctness.2 Current permissions and deletion policy apply at every retrieval, source-fetch, citation-preview, cache, and rollback path, even if an older corpus snapshot allowed access.3 Full replay is possible only within the permitted retention window and with compatible retained inputs.
The latency objective also has to be divided among components. The grounded-answer path runs several steps before the model produces its first token, so retrieval, reranking, and prefill share one TTFT budget. A deterministic planning allocation is an intentionally chosen time allowance for one component in that design budget. It is not a percentile of the end-to-end latency distribution. Per-stage p95 observations can inform those allocations, but their sum does not determine end-to-end p95. The team compares end-to-end TTFT measurements under the expected workload with the p95 objective, or validates a statistical model of the joint stage latencies, before deciding whether the objective is met.
Example: time to first token for one grounded answer
The service objective is p95 TTFT \(\leq\) 1,200 ms. Measured per-stage p95 values inform the following chosen planning allocations:
- Before retrieval: Gateway and authorization 20 ms, query embedding 15 ms.
- Retrieval (Chapter 5): Filtered dense and sparse search 60 ms, reranking 40 candidates 120 ms, packing 5 ms.
- Engine (Chapter 4): Queue 150 ms, prefill of a 6,500-token prompt 600 ms.
- Planned total: The chosen allocations sum to 20 + 15 + 60 + 120 + 5 + 150 + 600 = 970 ms, leaving 230 ms of allocation headroom.
- Option A: Reranking 80 candidates instead of 40 adds about 120 ms, for a 1,090 ms allocation. It fits this allocation, with 110 ms of allocation headroom.
- Option B: An answer-groundedness judge requires the generated answer. Holding the stream for that judgment adds answer-generation time plus about 900 ms for the judge. Even if answer-generation time is omitted, the planning allocation is already about 1,870 ms, above the objective. The judge therefore runs on sampled traffic after the answer is streamed.
- Conclusion: Deeper retrieval and inline evaluation compete for the same design allocation. The team compares end-to-end TTFT measurements under the expected workload with the p95 objective to decide whether either request path meets it.
Release observation window: collecting enough evidence before promotion
A candidate is promoted only after its evidence covers the traffic patterns that matter:
- Baseline: Compare a candidate with the last known-good release on the same request mix and protected quality slices.
- Window: Collect enough shadow or canary traffic to include peak and off-peak behavior, cache warm-up, retries, delayed evaluator labels, and one planned rollback rehearsal.
- Measures: Track availability, TTFT and ITL, retrieval recall and freshness, groundedness, authorization failures, cost, and evaluator disagreement with release identity attached.
- Decision: Promote only when every required slice and guardrail has data for the full window. Otherwise, extend observation, repair the candidate, or keep the known-good release.4
- Conclusion: A defined measurement window prevents the team from accepting a short favorable sample as complete release evidence.
8.2 Coordinating artifact versions and release responsibilities
The product guarantees are explicit. Incidents become untraceable when ‘the model,’ ‘the data,’ or ‘the index’ refers to a mutable alias instead of a specific generation. The release records which run built each artifact and which later request or deployment consumed it.5
| Artifact | Example immutable identity | Owning layer | Required references |
|---|---|---|---|
| Source corpus | corpus_manifest=sha256:... |
Data governance/storage | Source object versions and ACL snapshot |
| Chunk set | chunker:v4/corpus:41 |
Retrieval pipeline | Corpus, code, parameters |
| Embedding set | embed:e3/chunks:41 |
Retrieval pipeline | Chunks, model revision, normalization |
| ANN collection | collection:c41-e3-i7 |
Vector platform | Embedding set and index config |
| Training run | experiment:ft-2026-071 |
Training/experiment tracking | Dataset, base model, code, config, image |
| Model checkpoint | model:answerer-v12 |
Model registry | Training run and checkpoint manifest |
| Serving bundle | bundle:b12.4 |
Inference platform | Model, tokenizer, engine, quantization, prompt defaults |
| Evaluation report | eval:grounded-v9/b12.4/c41 |
Evaluation platform | Suite, judge, bundle, collection |
| Deployment | release:r218 |
Control plane | Bundle, resolved collection and query encoder, report, approval |
| Request trace | trace_id + release:r218 |
Observability | Runtime spans, retrieved IDs, output digest |
The release descriptor introduced in Section 5.5 now identifies the serving bundle, retrieval collection, query encoder, prompt and policy configuration, and evaluation evidence used together for one request or run. A candidate can have a descriptor before approval. The descriptor identifies the combination, while the release decision determines whether it may receive traffic.
Artifact governance: immutable resolution of routing aliases
Separating mutable routing aliases from immutable artifacts preserves experiment reproducibility:
- Human-friendly aliases such as
production,current-corpus, orlatest-modelare mutable routing pointers and therefore unsuitable as unresolved experiment inputs.6 - At entry, a run or request resolves one immutable release descriptor and carries that same bundle, collection, query encoder, and policy-compatible configuration across retries. In-flight requests continue using their resolved combination unless current policy requires them to stop.
For one failed request, the trace identifies the deployment. The deployment identifies the bundle, collection, and approved evaluation, which in turn identify their build runs and immutable inputs. In the release direction, a candidate receives a deployment ID only after every required upstream identity is present.
Release governance: responsibilities and emergency escalation
Defining approval roles and escalation authorities prevents unvetted production changes:
- Producing teams are responsible for the completeness and evidence of the artifacts they produce. The release owner assembles the compatibility record, and the operations owner manages rollout, rollback, and the evidence window.
- Approval records the candidate, evidence links, accountable person, and expiry or review point. Escalation identifies who can stop traffic when a hard guardrail fails. It is not replaced by a favorable average score.
8.3 Tracing documents from ingestion to retrieval
Those recorded links define the required connections. The first active path turns governed documents into searchable state without losing source, access, deletion, or version information. Committed generations and quality checks make the ingestion path reproducible.
The document ingestion path progresses through eight verifiable stages.
- Capture the source. A source snapshot records the document versions, access metadata, and deletion watermark in an immutable corpus manifest. This is the input that later checks compare against.
- Validate the intake. Validation checks file types, character decoding, duplicate rate, policy compliance, and item counts. An invalid batch stops here rather than producing searchable content.
- Create traceable chunks. A versioned splitter records each chunk’s content hash, access metadata, and a locator that can return a reader to the source. PDF pages, OCR regions, and text spans can need different locators. Byte offsets are used only when an extraction and normalization mapping has been tested.
- Produce compatible vectors. The embedding task records the model revision and normalization rule with the vectors. A later query must use the compatible query representation.
- Build a candidate index. The index task creates an approximate-nearest-neighbor (ANN) structure in a staging collection. Live traffic continues to use the current production alias.
- Reconcile the candidate. A coordinator compares vector counts, unique IDs, deletions, and partition coverage with the corpus manifest. A missing partition or an unprocessed deletion leaves the candidate incomplete.
- Evaluate the result. Benchmark retrieval cases run under expected load to check recall, latency, and access-filter behavior.
- Publish one committed generation. Only the coordinator records a durable committed-generation manifest and makes the candidate eligible for canary traffic. A successful worker process alone does not make the collection visible.
The following checks define completion for the candidate generation instead of relying on process exit status.7
| Gate | Pass criterion | Why exit code is insufficient |
|---|---|---|
| Completeness | Expected source/chunk/vector counts reconcile by tenant and partition | A worker can skip a partition and still exit successfully |
| Integrity | Hashes, dimensions, metadata schema, and unique IDs match the schema definition | Serialization may succeed with wrong shape or duplicates |
| Authorization | ACL mutation and negative-access tests pass | Readable vectors can still expose forbidden content |
| Retrieval quality | Required recall and citation coverage pass on held-out slices | An index can be healthy but semantically poor |
| Freshness | Inserted and updated content becomes searchable, and deleted content stops being returned, within the objective | Static benchmark ignores lifecycle lag |
| Recovery | Candidate restores or rebuilds within RTO | Replication status does not prove usable recovery |
The completion manifest is the commit record. Expensive workers may retry and write attempt-scoped shards, but only the coordinator can convert those outputs into a consumable generation after reconciliation. The promotion task reads that committed evidence instead of inferring completeness from worker status alone.
The candidate must remain immutable between checking and use. The conditional publication rules from Sections 3.4 and 7.6 reject old coordinators and delayed attempts. Collection-alias atomicity alone does not make a serving, retrieval, prompt, and cache change one transaction.8
8.4 Reproducible model experiments
The retrieval collection can now be reproduced. The generator or judge model may still need fine-tuning on examples from the company’s own documents, for example to follow the citation format reliably. Such a run can train a LoRA adapter (Section 4.6) or all weights, and when it is long and distributed it must survive interruption without losing experiment identity. One run manifest connects distributed strategy, scheduling, checkpointing, and experiment evidence.
Model-state and activation memory decide whether DDP, FSDP, tensor, pipeline, sequence/context, or expert parallelism, with or without activation recomputation, fits the constraint (Section 2.4). A LoRA run holds optimizer state only for the adapter, so it often fits where full fine-tuning would need sharding. The communication-heavy dimension uses the fastest available fabric, and the scheduler admits the complete topology. Each rank records world, node, local-rank, device, and shard identity.
| Stage | Required identity/evidence | Failure point |
|---|---|---|
| Allocation | Code commit, image digest, data manifest, base model, config, topology | Rendezvous and gang allocation |
| Active training | Loss, throughput, MFU, memory, power, communication, data time | Rank failure, straggler, NaN, device error |
| Checkpoint | All model/optimizer/RNG/sharding state and completion manifest | Partial save or lost node-local state |
| Resume | Checkpoint generation plus compatible code/config/topology rules | Silent restart from wrong step or sampler position |
| Held-out evaluation | Held-out suite, raw outputs, evaluator identity, quality/cost slices | Metric-only success without failed cases |
| Register | Validated immutable model version and lineage | Mutable path or incomplete bundle |
The distributed checkpoint saves the state supplied by the training application, subject to format and version compatibility.9 The experiment tracker compares candidate runs, while the orchestrator records task execution and retry history. The checkpoint manifest links both. Model registration occurs in a separate gate after held-out evaluation.
Reproducibility audit: identifying drift between twin execution runs
Comparing execution manifests reveals configuration drift and non-deterministic environmental factors:
- Same baseline: Both runs use the same dataset snapshot, evaluation suite, tokenizer, and serving test mix.
- Changed factors: Run B changes the sharding strategy and image digest. The manifest records those differences explicitly.
- Comparison: Run B improves throughput but increases recovery time and shows a quality drop on the long-context slice.
- Decision: The candidate is not promoted until the long-context regression is explained or repaired. The aggregate throughput gain alone is insufficient.
- Conclusion: Comparability comes from holding the declared inputs constant and making every changed factor visible.
8.5 Releasing serving and retrieval changes together
Model and collection artifacts are independently versioned. A grounded-answer response depends on their compatibility with the prompt, citation formatter, tokenizer, and evaluator, so promoting only one alias can create an untested combination. The deployable unit is one immutable release descriptor whose component combination has passed the linked artifact-pairing checks in Table 8.2, the staged evidence in Table 8.5, and the compatible encoder and collection relation in Section 5.5.
- Release candidate: One immutable serving bundle, retrieval collection, prompt/policy configuration, and evaluation report proposed for production traffic.
Each test stage adds realism and risk after the candidate has passed the preceding evidence check.
| Test stage | Traffic | Checks | Promotion condition |
|---|---|---|---|
| Interface | Synthetic | Startup, API schema, tokenizer, tool and citation parsing | All deterministic tests pass |
| Fixed benchmark | Recorded prompts and lengths | TTFT, ITL, throughput, memory, output format | No required regression |
| Offline quality | Held-out cases | Groundedness, citations, safety, retrieval quality | Slice thresholds and uncertainty pass |
| Shadow | Copied live inputs, no user-visible result | Load mix, routing, trace completeness, disagreement | No unexplained severe divergence |
| Canary | Small controlled live share | SLOs, errors, evaluator proxies, complaints, cost | Evidence window and guardrails pass |
| Production | Ramped share | Same guardrails plus capacity and rollback readiness | Stable after each ramp step |
The gateway stamps the resolved release, bundle, collection, prompt, and policy identities into request context. The selected router policy uses capacity and reusable KV prefixes, subject to model/configuration compatibility and the allowed tenant boundary.10 The engine records queue, prefill, decode, token, memory, and error evidence. Retrieval spans record filters, candidate stages, selected source IDs, and score conventions.
8.6 Linking request traces to population quality
The release can receive production traffic. Operators need both aggregate alerts and one replayable request path, while evaluators need enough governed context to judge behavior. Service telemetry, AI traces, sampled judgments, delayed labels, and release identities form one evidence path.11
| Evidence layer | Population signal | Request-level evidence |
|---|---|---|
| Gateway/router | Rate, errors, queue, retries, replica balance | Tenant class, release ID, route and retry spans |
| Retriever | Recall proxy, no-result rate, score and freshness drift | Query digest, filters, candidate IDs/scores, selected chunks |
| LLM engine | TTFT, ITL, throughput, KV pressure, batch size | Prompt/output token counts, replica, cache reuse, timing |
| Quality | Groundedness/citation/safety rates with uncertainty | Raw evaluator result, claim labels, parse status |
| Feedback | Complaint and acceptance rates by slice | Linked feedback type and timestamp |
| Cost | Tokens, GPU time, retrieval/rerank calls, judge sampling | Per-request attributed usage |
The telemetry owner applies the data policy when selecting trace samples and deciding which request fields remain available for replay. Aggregate metrics can remain longer than trace bodies, while trace bodies need a defined retention period and access rule.
Opaque IDs and hashes can still reveal sensitive information, especially when the original value is predictable. Redaction, encryption, access, and retention are implemented controls, not guarantees supplied by identifiers or telemetry conventions. When a payload expires, its identity may support lineage inspection but not full replay.12 Hashed or opaque source identities replace full content where possible, secrets and personal data are redacted before export, governed payloads are encrypted and access-controlled, and trace bodies expire independently of low-cardinality metrics. The evaluator revision and raw parse outcome help distinguish a judge change from a product change. The record also binds the prompt, rubric, model settings, preprocessing, and reference set. Controlled rescoring and human review of critical disagreements remain necessary.13
8.7 Finding the cause of a groundedness incident
The evidence path spans every component. The incident begins with an available service whose answers for newly added policy documents become less grounded. An ordered investigation localizes the failing artifact or stage before any component is blamed.
The following incident is a constructed example. Its numbers illustrate the diagnosis and recovery sequence, not results from an external production system.
Regression analysis: quantifying groundedness loss across document slices
Slicing benchmark evaluations pinpoints quality regressions introduced by updated retrieval collections:
- Production change: Release
r218keeps serving bundleb12.4but moves retrieval alias from collectionc40-e3-i7toc41-e3-i7. - Alert: The sampled groundedness rate falls from 0.91 to 0.78 for the
new-policyslice. HTTP errors and TTFT remain within objective. - Trace evidence: Failed answers have low retrieval scores and omit source IDs introduced in corpus generation 41.
- Lineage check: The corpus manifest expects 12.0 million chunks, but the committed collection contains 11.2 million. One tenant partition is absent.
- Pipeline evidence: An embedding task retried after a worker loss. Its completion wrapper reported process exit zero as success, while the release coordinator did not reconcile per-partition counts before promotion.
- Immediate mitigation: The release returns to the compatible
c40-e3-i7andb12.4combination only after current deletion and access checks pass, affected owners receive notification, and failed traces remain preserved. - Repaired candidate: A new candidate contains the rebuilt partition, reconciled counts and hashes, authorization/retrieval/freshness results, and shadow and canary evidence.
- Prevention: Completion-manifest reconciliation becomes a hard task gate, and compute tasks no longer change aliases.
- Conclusion: Availability and latency stayed normal while answer quality fell. Linked release, trace, storage, and pipeline records identified the incomplete retrieval collection.
| Hypothesis | Evidence that weakens or supports it |
|---|---|
| LLM model regressed | Weak: serving-bundle identity did not change, failures cluster by corpus slice. |
| Inference overloaded | Weak: queue, TTFT, ITL, memory, and error distributions remain stable. |
| Embedding model incompatible | Weak: both old and candidate collections use embedding version e3. |
| Access filter removed evidence | Possible until negative/positive ACL traces show correct filter decisions. |
| Index is incomplete | Strong: expected/actual count mismatch and missing partition IDs align with failed cases. |
| Judge changed | Weak: evaluator version is constant and user complaints corroborate the same slice. |
8.8 Regression prevention with evaluation gates
Rollback can restore the previous answer quality for requests covered by the known-good collection, provided that the older state still meets current access and deletion rules. Rebuilding one missing partition fixes this generation but leaves the pipeline capable of promoting another incomplete collection. A prevention control at the visibility point stops another invalid collection from reaching live traffic.
The prevention plan for this constructed incident uses seven connected controls. It does not claim that seven controls are necessary or sufficient for every system.
- Check the candidate before promotion. The verification coordinator checks partition row counts, unique IDs, vector dimensions, and access-control parity. This blocks an incomplete collection before it becomes visible.
- Keep writers away from the live alias. Compute workers write only to a candidate staging namespace. Only the verification coordinator has permission to change the live alias. Compute-worker identities cannot publish their partial output.
- Require a durable success record. The coordinator accepts a completion manifest tied to the reconciled candidate, not only a zero exit code. Hashes and signatures can add integrity checks when the threat model requires them, but they do not prove that the candidate is complete or authorized.
- Keep the failed case in the release suite. A targeted regression slice taken from the production incident gives the release gate a case that detects the same failure.
- Test recovery before an incident. A rehearsal removes a partition from a candidate and verifies that promotion stops without changing the live alias.
- Watch coverage after promotion. Searchable-count and partition-freshness signals reveal an unexpected loss of available content.
- Review canary evidence. Groundedness telemetry during the stabilization window checks whether the released system remains within its release objectives.
Release gate principle: optimal control placement
Placing preventive gates before deployment detects defects before user impact:
- The preventive gate belongs at the point where invalid state could become visible.
- A dashboard can detect a bad collection after promotion, while a commit manifest and promotion check block the invalid collection from live traffic.
8.9 Deciding release readiness from recorded evidence
The incident has produced a durable prevention control. Future releases need a compact cross-layer checklist so evidence is assembled before the rollout window rather than during an incident. Each responsible component contributes a verifiable artifact or measurement to the candidate release.
| Layer | Release evidence | Stop condition |
|---|---|---|
| Artifacts/lineage | Immutable model, tokenizer, image, prompt, data, collection, config, evaluator IDs | Any unresolved mutable dependency |
| Training (when used) | Run manifest, checkpoints, held-out quality, MFU/throughput, recovery test | Missing state or unexplained instability |
| Serving | Compatibility, load, soak, cold-start, capacity, rollback results | SLO or rollback failure |
| Observability | Correlated logs/metrics/traces, dashboards, alerts, redaction/retention | Blind request stage or ungoverned payload |
| Evaluation | Calibrated judge, raw results, consistency, slices, human adjudication | Unstable or unvalidated critical slice |
| Storage/retrieval | Completeness, integrity, ACL, freshness, recall, tail latency, restore | Count/hash/authorization mismatch |
| Pipeline | Idempotency, retry/backfill tests, task descriptions, promotion gate | Side effect can repeat or partial output can commit |
| Operations | Responsible person, runbook, canary plan, capacity, incident and rollback authority | No accountable decision path |
A stop condition has priority over a favorable aggregate score.
The release policy treats retrieved documents as untrusted evidence, not privileged instructions. Authorized content can still contain indirect prompt injection: instructions embedded in external content that try to redirect the application.14 A retrieved document can contain text that tries to override application instructions, disclose protected information, or influence a tool call. Authorization permits retrieval of the document. It does not authorize the model or application to obey text that conflicts with its governing instructions and tool rules. Adversarial retrieval and evaluator cases test refusal and containment where untrusted content enters the request. No single filtering or prompting rule is assumed to solve the attack class. High overall groundedness cannot compensate for an authorization failure, and high throughput cannot compensate for an untested rollback path. The release owner resolves each red condition explicitly: failed authorization, a missing required artifact, an unmet evaluation threshold, an incomplete rollback test, or an unresolved canary stop condition. Absence of evidence is not a pass.
Operations governance: approval policy versus escalation authority
Dividing policy approval from emergency escalation enables rapid incident response:
- Approval confirms that the recorded evidence satisfies the release policy and names the accountable release owner.
- Escalation covers a live stop: The on-call or incident authority can pause traffic, restore the known-good aliases, preserve evidence, and notify the affected owners while the release decision is revisited.
8.10 Improving the system from production feedback
The checklist can approve one candidate. Production data, traffic, providers, model aliases, document corpora, hardware, and costs continue to change after release. Each release is a measured hypothesis, and its failures become inputs to the next reproducible evaluation and pipeline run.
The operating loop is now complete. For example, a groundedness alert on newly added documents signals a change in measured groundedness for that document slice. A correlated trace supplies a replay case, the evaluation suite checks the repaired candidate, and the pipeline records the resulting serving and retrieval pair. Canary evidence then supports promotion or returns traffic to the known-good pair. The same evidence path supports that investigation. Fixed evaluator identity, controlled rescoring, and adjudication are needed to distinguish a product change from judge instability before the next release decision.
Development and incident-regression cases stay distinct from independently held-out evaluation. User acceptance is not automatically a correct label, and feedback can change future data in harmful ways. Improvement must be measured rather than assumed from completing the loop.15
| Question | Evidence source | Typical response |
|---|---|---|
| Is the service available? | SLO metrics, errors, traces, capacity | Serving-path operation and capacity adjustment |
| Is the behavior good? | Evaluators, human labels, complaints, business outcomes | Failure triage by slice |
| Why did it happen? | Trace-to-bundle-to-run-to-artifact lineage | Hypotheses for investigation and controlled replay |
| Can it be reproduced? | Permitted retained inputs, images, configs, seeds, and raw outputs | Replay within the retention and runtime compatibility boundary |
| Is the change better? | Held-out and production-like comparative evidence | Candidate approval, rejection, or revision |
| Can it be released safely? | Compatibility, canary, rollback, recovery, policy gates | Controlled promotion |
| Did the fix persist? | Post-release guardrails and incident slice | Closure after the evidence window |
Chapter 8 summary
Key system integration and operational principles established in this chapter:
- Core mechanisms: A release descriptor pins one tested serving bundle, retrieval collection, prompt, and policy. Request traces carry those identities, so a groundedness alert can be followed to a collection build and its pipeline run.
- Governing trade-offs: Retrieval depth, reranking, and inline evaluation compete for one TTFT budget. A longer observation window gives stronger release evidence but slows delivery.
- Failure modes & defenses: An incomplete collection passed because a worker’s exit code counted as success. Reconciling counts before promotion, keeping compute tasks away from the live alias, and adding the incident to the release suite stop the same failure before users see it.
Chapter checkpoint
Review Questions 39–41 in Appendix B, Section B.1, to test incident diagnosis, the placement of the prevention gate, and the end-to-end latency budget.
System synthesis: reliable AI platform engineering principles
Unified platform engineering principles connecting the complete machine learning lifecycle:
- In this platform design, reliable operation depends on explicit connections among versioned artifacts, compute strategies, serving stages, telemetry, evaluation, durable state, and orchestration.
- For every result, the operating model identifies its exact inputs, responsible component, validation evidence, failure behavior, reproduction path, and rollback path.
Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. Curran Associates. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html. The original RAG paper combines retrieval with generation in tested QA/generation tasks. Authorization, citation checking, and replay are additional book requirements, not guarantees established by that paper.↩︎
Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations. In Proceedings of EMNLP 2023 (pp. 6465–6488). Association for Computational Linguistics. DOI: 10.18653/v1/2023.emnlp-main.398. https://aclanthology.org/2023.emnlp-main.398/. ALCE measures correctness and citation quality separately and finds incomplete support in its tested systems. Source entailment does not establish that a source is true, current, or authorized.↩︎
Microsoft. (2026, August 24). Security filters for trimming results in Azure AI Search (Examples use API 2026-04-01). https://learn.microsoft.com/en-us/azure/search/search-security-trimming-for-azure-search. The official security-filter pattern requires every query to use the relevant permissions and is not itself authentication. The surrounding system also applies current-policy checks to derived caches and rollback.↩︎
Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data (pp. 1123–1132). IEEE. DOI: 10.1109/BigData.2017.8258038. https://storage.googleapis.com/gweb-research2023-media/pubtools/4156.pdf. The ML Test Score supports important-slice, pre-serving, canary, monitoring, and rollback checks. A full time window alone is not a sample-size or uncertainty guarantee.↩︎
Moreau, L., & Missier, P. (Eds.). (2013, April 30). PROV-DM: The PROV data model. W3C Recommendation. https://www.w3.org/TR/2013/REC-prov-dm-20130430/. PROV-DM represents entity/activity generation, use, and derivation. It does not verify the truth of links or prescribe this release workflow.↩︎
MLflow contributors. (n.d.). MLflow model registry (MLflow 2.4.2 archived documentation). https://www.mlflow.org/docs/2.4.2/model-registry.html. MLflow 2.4.2 explicitly permits alias reassignment. An alias can be accepted at entry if resolved once to a retained immutable version. Repeated resolution is the unsafe case.↩︎
Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781–1794. DOI: 10.14778/3229863.3229867. https://www.vldb.org/pvldb/vol11/p1781-schelter.pdf. The original data-quality system supports declared constraints and aggregation-based checks. Matching totals and hashes do not establish semantic completeness, correct source locators, or ACL propagation.↩︎
Qdrant. (n.d.). Collections (Collection aliases and switching). https://qdrant.tech/documentation/manage-data/collections/. Qdrant’s atomic grouped alias update is confined to its API. The broader request-pinning and coordinator protocol remains a book design.↩︎
PyTorch Contributors. (2025, June 16). Distributed Checkpoint: torch.distributed.checkpoint. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/distributed.checkpoint.html. The distributed-checkpoint API saves supplied state with format/version constraints. The surrounding workflow adds explicit sampler, scheduler, scaler, step, manifest, and registration requirements.↩︎
NVIDIA. (n.d.). Routing concepts (Dynamo 1.0.2 documentation). https://docs.nvidia.com/dynamo/v1.0.2/components/router/routing-concepts. Dynamo 1.0.2 gives one KV-aware routing mode with load/cache costs. Other routers and modes differ, and the citation is not an unconditional latency or isolation guarantee.↩︎
OpenTelemetry authors. (2026, January 14). Traces (OpenTelemetry documentation). https://opentelemetry.io/docs/concepts/signals/traces/. OTel supplies span context, attributes, and links. Joining releases and delayed judgments requires application instrumentation and does not by itself prove causation or full replayability.↩︎
OpenTelemetry authors. (2026, January 14). Handling sensitive data (OpenTelemetry security guidance). https://opentelemetry.io/docs/security/handling-sensitive-data/. OTel recommends minimization and redaction and warns that predictable inputs can be recovered from hashes. Retention, encryption, and authorization still need implementation and policy.↩︎
Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena (arXiv:2306.05685v4). arXiv. Author manuscript. The judge study documents bias and reasoning limits in preference evaluation. Stable identity alone does not make repeated judgments stable or establish a groundedness error rate.↩︎
Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection (arXiv:2302.12173v2). arXiv. Author manuscript. The original study demonstrates attacker instructions in external content affecting tested LLM-integrated applications. It does not establish that a particular defense completely solves the problem.↩︎
Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503–2511. https://papers.nips.cc/paper_files/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf. Hidden Technical Debt Section 4 explains direct and hidden feedback that changes future data. A feedback loop can worsen behavior rather than automatically improve it.↩︎