Part III: Automated workflows and reliable system operation

Independent endpoints, storage systems, and training jobs create coordination work. Automated dependencies, explicit retry behavior, and incident recovery provide one way to manage that work.

This part connects repeatable Airflow workflows and immutable run records to coordinated release checks and incident recovery (§§ 7.4, 7.6, 7.8, 8.5, 8.7).

Chapter 7 orchestrates reproducible ML pipelines, and Chapter 8 operates the complete grounded-answer system and closes the production-to-research improvement loop.

The map follows dependency-aware workflows into coordinated release evidence, incident diagnosis, and regression prevention.

A Part III map with Chapter 7 workflow definition, isolated workers, safe retry, and committed evidence, and Chapter 8 coordinated release, observation, diagnosis, restore or repair, and prevention, with an orange loop returning a regression case to the next run.
Figure 1: How Part III carries evidence between its chapters. Dependency-aware workflows produce committed run manifests, and the grounded-answer system turns an incident into a regression case for the next pipeline run.

Chapters in this part

  • Running repeatable ML workflows: Every week the policy corpus changes. An update can require document validation and new embeddings before the index is rebuilt. Evaluation then informs the release decision. If a step fails halfway or runs twice, an operator needs its recorded inputs and output state to decide how to recover. The same records let another operator continue the work safely.
  • Operating a grounded-answer system end to end: Suppose that after a routine weekly update, answers about newly added policies start citing the wrong documents or none at all, while latency and error rates stay normal. The groundedness defined in the Introduction has fallen even though the service is available. The preceding chapters produced versioned artifacts, recoverable distributed jobs, durable storage, serving bundles, a versioned retriever, evaluation evidence, and retryable pipelines. This chapter operates them as one system whose release evidence connects such an alert to a verified recovery.