Part III: Automated workflows and reliable system operation
Independent endpoints, storage systems, and training jobs create coordination work. Automated dependencies, explicit retry behavior, and incident recovery provide one way to manage that work.
This part connects repeatable Airflow workflows and immutable run records to coordinated release checks and incident recovery (§§ 7.4, 7.6, 7.8, 8.5, 8.7).
Chapter 7 orchestrates reproducible ML pipelines, and Chapter 8 operates the complete grounded-answer system and closes the production-to-research improvement loop.
The map follows dependency-aware workflows into coordinated release evidence, incident diagnosis, and regression prevention.
Chapters in this part
- Running repeatable ML workflows: Every week the policy corpus changes. An update can require document validation and new embeddings before the index is rebuilt. Evaluation then informs the release decision. If a step fails halfway or runs twice, an operator needs its recorded inputs and output state to decide how to recover. The same records let another operator continue the work safely.
- Operating a grounded-answer system end to end: Suppose that after a routine weekly update, answers about newly added policies start citing the wrong documents or none at all, while latency and error rates stay normal. The groundedness defined in the Introduction has fallen even though the service is available. The preceding chapters produced versioned artifacts, recoverable distributed jobs, durable storage, serving bundles, a versioned retriever, evaluation evidence, and retryable pipelines. This chapter operates them as one system whose release evidence connects such an alert to a verified recovery.