7  Running repeatable ML workflows

Every week the policy corpus changes. An update can require document validation and new embeddings before the index is rebuilt. Evaluation then informs the release decision. If a step fails halfway or runs twice, an operator needs its recorded inputs and output state to decide how to recover. The same records let another operator continue the work safely.

Selective recovery requires a record of which work was ready, which inputs it used, and which output was committed. Dependency graphs organize tasks with explicit inputs, outputs, resources, retries, and evidence.

Chapter map: repeatable ML workflows

The sections answer four linked questions:

  • 7.1–7.3: What does an orchestrator add to a script, which tasks does a dependency graph allow to run, and what must each task declare?
  • 7.4–7.5: Which Airflow components parse, schedule, launch, and record a run, and how does a DAG launch validation, embedding, and evaluation in isolated pods?
  • 7.6–7.7: How can a task be retried without duplicating or partially replacing results, and how are common task patterns reused?
  • 7.8–7.9: Which system holds which record of a run, and how do teams share one cluster without reading each other’s data?
Nine numbered panels show manual commands becoming a dependency-aware DAG, immutable task contracts, a separated Airflow control plane and task execution boundary, a complete ML workflow, safe retry with checked completion, reusable templates, run manifests, and separate team controls.
Figure 7.1: Dependencies, immutable task contracts, separated orchestration and execution, safe retry, reusable templates, run records, and tenant controls make workflow runs repeatable.

7.1 From manual scripts to automated pipelines

A single script can run ingestion, training, evaluation, and deployment in sequence. As the workflow grows, bespoke scripts must implement durable coordination for parallel branches, selective retries, backfills, resource placement, responsibility, and persistent run state. The workflow system supplies reusable mechanisms for that work rather than making it impossible in one process. Dependency definitions stay separate from execution, and task-attempt state persists across failures.

  • Pipeline: A reproducible set of tasks and dependencies that transforms versioned inputs into versioned outputs while recording execution state.

  • Orchestrator: A system that coordinates when and where work runs and records its execution state. Workflow orchestration also checks dependencies between tasks and records their attempts.

The cluster orchestrator in Chapter 1 places workloads on machines. A workflow orchestrator coordinates the steps of a larger job. Both can participate in one pipeline: Airflow checks whether a task is ready, then a pod operator asks Kubernetes to launch its workload.1

Orchestration is not the model code. It decides when a task is eligible, where it runs, how the system records its state, and what happens after failure. The task implementation performs domain work such as preprocessing or training. The artifact store stores large outputs. The metadata store stores orchestration state and small references.2

Table 7.1: Configured orchestration goals around task code, not automatic guarantees of every pipeline product.
Concern Ordered script Orchestrated pipeline
Failure Restart manually or add ad-hoc checkpoints Retry or resume the failed task under policy
Parallelism Manual process management Independent branches run concurrently
History Logs scattered by process Persistent runs, attempts, owners, and timestamps
Scheduling External cron or operator Calendars, events, backfills, and concurrency limits
Resources Inherited host environment Per-task image, CPU, memory, GPU, queue, and secrets
Lineage Implicit file names Versioned input/output references and run identity

7.2 Scheduling task dependencies with directed acyclic graphs

The workflow separates responsibilities. The orchestrator still needs an unambiguous rule for which tasks may run and which work can happen in parallel. Tasks are nodes, and prerequisite relationships form directed edges without cycles.

  • Directed Acyclic Graph (DAG): A graph whose directed edges encode dependency order and whose lack of cycles means that a topological execution order exists.3
A validated dataset splits into parallel embedding-index and model-training branches, then both converge on a promotion gate.
Figure 7.2: The DAG exposes two parallel branches that must both reach the promotion gate; edges order work, while artifact readiness still needs an explicit check.

In this data contract, A → B means B requires A’s validated output, and A does not invoke B directly. Airflow edges instead govern eligibility through task state and trigger rules, so artifact readiness needs an explicit check. Independent branches can execute concurrently. A cycle has no valid strict dependency order and is normally rejected by a DAG system. Feedback over time belongs in separate runs or external events.

\[ T_{\mathrm{DAG}} \ge \max_{p \in P} \sum_{i \in p} t_i \tag{7.1}\]

\(T_{\mathrm{DAG}}\) is the completion time of one run, \(P\) is the set of paths from a starting task to a final task, and \(t_i\) is the duration of task \(i\). The longest path, the critical path, sets the lower bound because its tasks must run one after another.

Example: scheduling duration in a branched DAG

For a finite DAG with fixed task durations and finish-before-start dependencies, the longest path is a completion-time lower bound. Equality requires enough resources and no omitted queue, transfer, or retry delays:4

  • Shared start: Validation takes 4 minutes.
  • Training branch: Training takes 38 minutes and model evaluation takes 8 minutes.
  • Retrieval branch: Embedding and index building together take 20 minutes, and index validation takes 6 minutes.
  • Shared release: The promotion gate takes 3 minutes after both branches finish.
  • Path sums: Adding the task durations gives 53 minutes for the training path and 33 minutes for the retrieval path.
  • Conclusion: The 53-minute training path controls the lower bound. Accelerating the 33-minute branch alone cannot reduce end-to-end time below 53 minutes.

Workflow analysis: critical path scheduling implications

Identifying workflow bottleneck dependencies directs pipeline optimization efforts:

  • Doubling workers does not accelerate a strictly serial chain. With fixed dependencies and durations, shortening the longest paths reduces the lower bound. Tied critical paths may need joint improvement, and changed resource waits or task internals require a new calculation.
  • A slow off-path task can still consume capacity and delay later runs, so capacity and critical-path analysis remain separate.

7.3 Defining immutable task inputs and outputs

The DAG states dependency order. A box named ‘train’ is still too vague to retry safely or reproduce after an image, dataset, or configuration changes. Every task has immutable inputs, explicit outputs, resource needs, failure behavior, and validation.

  • Task specification: The complete contract for one pipeline task: identity, immutable inputs, outputs, runtime, resources, retry behavior, validation, and side effects.

  • Idempotent task: A task whose repeated execution with the same declared inputs reaches the same externally visible result without corrupting state or duplicating effects.5

Declared task fields let both the scheduler and the next task check which inputs and outputs belong to this run.

Table 7.2: A task description makes retry, audit, responsibility, and downstream consumption explicit.
Task field Example Why it matters
Identity train_model:v7 plus code commit Connects implementation to a run
Inputs Dataset manifest, base checkpoint, config hash Prevents dependence on mutable ‘latest’ paths
Outputs Checkpoint URI, metrics URI, completion manifest Makes downstream dependencies inspectable
Runtime Immutable image and entry point Identifies image-contained dependencies; record host and external dependencies separately
Resources 8 GPUs, 96 CPU, 1 TB memory, fabric constraint Enables placement and capacity decisions
Retry policy Two retries for temporary infrastructure failure, no retry on schema error Separates recoverable from deterministic failure
Validation Schema, counts, hashes, metric bounds Prevents success from meaning only exit code zero
Side effects Registry alias update or notification Requires deduplication, transaction, or manual gate

Large objects bypass the orchestration metadata database and inter-task message fields. The artifact store holds them, while tasks pass a small immutable reference plus integrity metadata.6 Each task writes to a unique attempt or generation path, validates its output, and conditionally publishes a completion pointer. Only then does it report task success. Section 7.6 applies the conditional-publication rules of Section 3.4 to task retries.

7.4 Coordinating workflow runs with Airflow

Task descriptions explain what each DAG node does. A concrete orchestrator still needs components that parse definitions, create runs, choose ready tasks, launch tasks, and expose state. Airflow’s components divide those responsibilities around the DAG file.

Control services contain the DAG processor, orchestration metadata, scheduler, Executor, UI and API server. The Executor launches the isolated supervised runtime, where the Task SDK supervisor starts user task code and reports state through the Execution API. User code exchanges business data only with external input and output storage.
Figure 7.3: The scheduler’s Executor launches a supervised task runtime. The Task SDK supervisor exchanges task state through the Execution API, while user task code reads and writes business data without direct metadata-database access.

In Airflow 3.3.1, the separate DAG processor parses definitions, the scheduler contains the executor, and a triggerer is needed when deferrable operators are used. The API server supports the user interface and REST API as well as the Execution API used to report task state. In a supervised Python Task SDK run, the supervisor communicates with that Execution API on the task’s behalf. Task code does not directly access the metadata database. Earlier Airflow layouts can differ from these documented roles.7

Table 7.3: Airflow 3.3.1 component responsibilities. Earlier versions can use different process layouts.
Airflow component Ownership
DAG processor Imports Python definitions and serializes valid DAG structure
Scheduler Creates scheduled runs, evaluates dependencies, and queues eligible task instances
Executor Within the scheduler, dispatches tasks to the selected local or remote execution backend
Workers / task pods Task execution in the selected environment with reported state
Metadata database Persists run, task, attempt, and scheduling state; secrets can use external backends
API server Serves the UI and REST API, and the Execution API used by supervised task runtimes to communicate state
Artifact/log backends Task logs and large domain outputs outside scheduler state
Triggerer (when used) Runs triggers for supported deferrable operators

The UI is a view of persisted state, not the execution engine. A success state can follow normal operator completion or manual marking, so it does not prove that domain checks ran or that the model is good. DAG-run status also depends on leaf states and trigger rules: a permissive successful leaf can mask an earlier failure. Release gates therefore consume explicit validation evidence, not merely a green run.8 Domain validation belongs inside task descriptions and release gates.

7.5 Running complete ML workflows in Airflow

The control-plane roles and task description are explicit. The Python definition describes dependencies without performing expensive work during import. Lightweight DAG parsing leaves domain code to isolated task runtimes.

Code example: Airflow dependency/runtime skeleton with external constants. No explicit retries or complete publication protocol are shown; KPO does not require KubernetesExecutor.

from airflow import DAG
from airflow.providers.cncf.kubernetes.operators.pod import (
    KubernetesPodOperator,
)
from datetime import datetime

with DAG(
    dag_id="versioned_rag_release",
    start_date=datetime(2026, 1, 1),
    schedule=None,
    catchup=False,
    max_active_runs=1,
    tags=["mlops", "rag"],
) as dag:
    validate = KubernetesPodOperator(
        task_id="validate_sources",
        image=PIPELINE_IMAGE,
        cmds=["python", "-m", "pipeline.validate"],
        arguments=["--manifest", DATASET_MANIFEST],
    )
    embed = KubernetesPodOperator(
        task_id="build_embeddings",
        image=PIPELINE_IMAGE,
        cmds=["python", "-m", "pipeline.embed"],
        arguments=[
            "--manifest", DATASET_MANIFEST,
            "--embedding-version", EMBEDDING_VERSION,
        ],
    )
    evaluate = KubernetesPodOperator(
        task_id="evaluate_retrieval",
        image=PIPELINE_IMAGE,
        cmds=["python", "-m", "pipeline.evaluate"],
        arguments=["--suite", EVAL_SUITE, "--candidate", COLLECTION_ID],
    )

    validate >> embed >> evaluate

Code walkthrough: Airflow dependency and runtime skeleton

This skeleton defines task order and isolated runtimes. Later task contracts provide the output checks, retry rules, and publication steps needed for safe recovery:

  • DAG definition: The DAG object stores scheduling and concurrency policy. schedule=None disables automatic time scheduling, while runs can be triggered through the UI, API, command-line tool, or other control paths.
  • Concurrency guard: max_active_runs=1 limits overlapping runs of this DAG. It is not a lock against other DAGs, operators, or delayed external actions changing the production alias.9
  • Task runtime: Each KubernetesPodOperator launches a pod for domain work after DAG import and does not require KubernetesExecutor. PIPELINE_IMAGE must actually be digest-pinned, and host/runtime dependencies remain separate.10
  • Artifact lineage: Task arguments carry dataset, embedding, evaluation-suite, and candidate-collection identities.
  • Dependency state: The shift operators create validation → embedding → evaluation edges. Airflow records each task instance and attempt.
  • Result: The scheduler owns eligibility and task state, while isolated containers own validation, embedding, and evaluation work.
  • Limits: The uppercase constants and domain programs are supplied externally. This snippet has no explicit retry policy. A production DAG still needs explicit output manifests, retries by failure class, secrets and service accounts, resource requests, timeouts, alerts, and a separate gated promotion task.

Orchestration principle: DAG import performance safeguards

Maintaining lightweight DAG definitions protects orchestrator scheduler performance:

  • Import-time dataset downloads, GPU initialization, external API calls, and large namespace scans slow parsing and can make scheduler health depend on external systems.
  • The DAG processor repeatedly parses definitions. Expensive import-time work can delay parsing and scheduling visibility, while one failed import does not imply that every DAG is broken.11

7.6 Safe re-execution and idempotent tasks

The example DAG can launch isolated tasks. Retries and historical runs will repeat code, so ambiguous partitions and irreversible side effects can duplicate or overwrite production state. Each run is bound to a logical data interval or immutable release identity, while computation and promotion remain separate.

Table 7.4: Proposed retry patterns. Publication additionally needs conditional winner selection, stale-attempt rejection, and recovery after uncertain acknowledgement.
Mechanism Safe pattern Failure to avoid
Retry Attempt-scoped output, validation, then one committed generation Appending duplicate rows or partially replacing a model
Backfill The run’s logical interval and versioned upstream snapshot Reading today’s mutable input for an old date
Catchup Enable only when every historical interval is meaningful and resourced Accidental flood of years of runs
Sensor / wait Supported deferrable operator plus triggerer and timeout; pool policy remains configurable Occupying a worker while polling indefinitely
External API Idempotency key with persisted response identity Duplicate billing, notification, or deployment
Promotion Policy gate and conditional shared-store update after validation A retried compute task silently deploys

Retry policies distinguish temporary infrastructure failures from code or data failures that will recur with the same inputs. Exponential backoff multiplies the delay between successive attempts, usually up to a configured limit. This delay gives a temporarily unavailable dependency time to recover. It does not fix a schema mismatch. After the retry budget, the task state, relevant logs, input identities, and responsible component remain visible without reconstructing the run from a terminal.

Idempotency walkthrough: safe partition task retry without duplicate writes

The protocol uses staged immutable output and conditional publication, not a formal two-phase commit protocol:

  • Attempt 1: Partition 07 writes to an attempt-scoped path, loses its worker, and remains uncommitted.
  • Eligibility: The configured retry policy or task code classifies the failure as retryable and preserves the same logical interval and input manifest. The scheduler does not infer arbitrary domain failure classes.12
  • Attempt 2: The task writes a new generation, validates row counts and hashes, and emits one completion record keyed by task, interval, and input identity.
  • Commit: A coordinator applies the conditional-publication rule of Section 3.4 with a key made of task, interval, and input identity. A conditional create selects one immutable winner, a lost response is resolved by reading the winner record back before any retry, and a late attempt cannot overwrite the winner.13
  • Conclusion: Attempt-scoped outputs and one publication step keep retries from duplicating or partially replacing state.

7.7 Building reusable task templates

Retry behavior and artifact publication are now safe. Teams still lose consistency when every DAG reimplements data validation, distributed training, inference, evaluation, and registration differently. Parameterized task blocks package common responsibilities through stable inputs, outputs, checks, and platform defaults.14

For example, a reusable validation block accepts an input-manifest ID, a digest-pinned image, a resource request, and a retry class. It validates the input, writes an output manifest, records telemetry, and passes its publication check before a downstream task can use the result. The task-specific part is the domain validation rule and its thresholds. The platform-owned part is identity, secret delivery, scheduling, retry recording, telemetry, and conditional publication.

Input manifest, image digest, resources and retry policy feed a reusable validation task. Platform identity and scheduling controls surround it. Validation evidence and conditional publication gate downstream work.
Figure 7.4: A reusable validation task combines versioned inputs with platform controls. Its domain check produces an output manifest and telemetry; only validated output proceeds through conditional publication. Identifier and resource values are illustrative.

Each reusable block has declared inputs, committed outputs, and validation checks.15

Table 7.5: Reusable blocks give domain code and platform control a consistent interface.
Block Consumes Produces Primary gate
Data validation Source snapshot and schema Validated manifest and profile Schema, counts, drift, policy
Training Dataset, base model, config, image Checkpoint and training evidence Finite loss, throughput, checkpoint integrity
Batch/online inference Serving bundle and input set Outputs and traces Interface, latency, errors
Evaluation Outputs, references, evaluator version Metrics, slices, failed cases Quality/safety thresholds
Registry Validated artifact and metadata Immutable version and lineage Completeness; integrity and authenticity checks required by policy
Deployment Approved serving bundle Candidate endpoint and rollout state Health, canary, rollback

In the platform contract, a reusable block should standardize identity, evidence, secrets, resource requests, telemetry, and output publication while allowing task-specific domain parameters. This division lets a model team change architecture without reimplementing cluster scheduling or artifact lineage.

7.8 Recording artifact evidence in run manifests

The pipeline can produce versioned artifacts through reusable blocks. Five systems now record overlapping facts, and copying all data into every system obscures the authoritative record. Each evidence class has one primary system, and stable identifiers connect the records.

Table 7.6: Primary responsibilities. Pending-link records and reconciliation handle partial cross-system writes.
System Canonical responsibility Cross-link
Orchestrator DAG run, task attempts, dependency state, schedule, retries Experiment run ID and artifact URIs
Experiment tracker Parameters, metrics, model/evaluator versions, comparable research runs DAG/task ID and deployment candidate
Artifact/registry store Immutable files, manifests, versions, promotion aliases Producer run and consumer release
Observability backend Service logs, metrics, traces, alerts, incidents Serving bundle, DAG run, experiment and request IDs
Evaluation store Versioned cases, references, raw judgments, slice reports Experiment, bundle, dataset and evaluator IDs

A pipeline task can start an experiment run and emit its ID. The tracker records comparable parameters and metrics. The task publishes artifact URIs and commits its status. A later deployment writes the serving-bundle ID into traces. The same identifiers let an incident move backward from one request to its bundle, evaluation report, experiment, DAG run, and source artifacts.

Architecture boundary: authority mapping across systems

The responsibility map reduces ambiguity. It does not make cross-system writes atomic:

  • The orchestrator is authoritative for task attempts and dependency state, the experiment tracker is authoritative for comparable parameters and metrics, the artifact registry is authoritative for immutable files and promotion aliases, telemetry is authoritative for runtime observations, the evaluation store is authoritative for cases and raw judgments.
  • Cross-links join these records without copying one system’s mutable state into another. A run manifest records the identifiers and the producer or consumer relationship.16

If artifact publication succeeds but a tracker update fails, a durable pending-link record keeps the missing update visible. A reconciliation task retries that link and checks the selected generation. The authoritative record and any delayed mirror remain distinguishable during recovery.

7.9 Isolating teams on shared research clusters

One pipeline has clear responsibilities and evidence. A research organization adds concurrent users, shared clusters, queues, secrets, quotas, interactive work, long-running training, and repeated evaluation. Code, compute, storage, observability, orchestration, and experiment systems collaborate while retaining distinct roles.

Each team has a service account and RBAC, a scheduler placement into compute, and a running container. Storage requests cross an IAM or ACL gate to the team's authorized storage area. Cross-tenant access is denied by configured controls, and access events produce protected audit records.
Figure 7.5: Two teams share scheduling and object storage while their configured identity and storage controls limit each workload to its authorized area. The repeated scheduler icon represents one service. The right-hand telemetry records describe storage-access audit evidence, not every workload log. Namespaces alone do not supply network, storage, or API isolation.

A shared research platform combines Git-hosted code, containerized workloads, Kubernetes scheduling, shared and object storage, telemetry, pipeline control, and experiment tracking. Interactive work and scheduled jobs can use the same immutable images and artifact identities, allowing an exploratory result to become a repeatable run.

Operators and researchers use different evidence from the same run. Pod state and task logs help operators investigate stalled or failed work. Training progress and evaluation help researchers decide whether the run is worth continuing. The same run ID connects both views.

Tenant boundary

  • In this shared cluster, namespaces or projects group each tenant’s resources, while quotas and queue limits bound resource use. Each workload runs under an identity with permission to call only the required APIs. Enforced network policies limit connections, storage access policies control reads and writes, and trace-store permissions control who can inspect recorded requests. A storage prefix groups objects but does not itself enforce access. Mutually untrusted code can require stronger node, sandbox, or cluster separation.17
  • Run and artifact identities remain globally traceable, while payload access follows tenant policy. A scheduler quota alone does not prevent one team from reading another team’s checkpoints or traces.18

Chapter 7 summary

Key orchestration and workflow principles established in this chapter:

  • Core mechanisms: A DAG states which tasks depend on which outputs, and its critical path sets an ideal lower bound on run time under the stated DAG assumptions. Each task declares immutable inputs, outputs, resources, retries, and validation. Airflow parses definitions, schedules ready tasks, and records attempts, while domain work runs in isolated pods. Run manifests link task attempts to experiments, artifacts, and traces.
  • Governing trade-offs: More workers cannot overlap tasks whose dependencies require them to run in sequence. They can reduce resource waits or shorten a task that can use more workers, so the critical path must be recalculated when task durations change. Reusable task blocks standardize identity, evidence, and publication at the cost of a shared platform contract. A shared cluster can improve utilization but needs separate access controls to keep tenants apart.
  • Failure modes & defenses: A green run can hide a skipped check, so release gates read validation evidence. Retries write attempt-scoped outputs and publish one winner. A scheduler quota does not stop one team from reading another’s checkpoints, so storage and trace permissions are separate controls.

Chapter checkpoint

Review Questions 35–38 in Appendix B, Section B.1, to test workflow records, idempotent retries, the separation of computation from production-alias promotion, and the critical-path calculation before examining the complete operational case.

Carry-forward result

The system now expresses work as a DAG of idempotent task descriptions, stores artifacts outside control-plane metadata, joins orchestration with experiments and telemetry, and supports selective retry and recovery.

The final chapter applies every layer to the grounded-answer service and follows a constructed quality regression from detection to verified release.


  1. Apache Software Foundation. (n.d.). KubernetesPodOperator (Kubernetes provider 10.22.0 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow-providers-cncf-kubernetes/stable/operators.html. KubernetesPodOperator launches pods through the Kubernetes API and does not require KubernetesExecutor. An image variable alone does not prove digest pinning, retries, or idempotent side effects.↩︎

  2. Baylor, D., Breck, E., Cheng, H.-T., Fiedel, N., Foo, C. Y., Haque, Z., Haykal, S., Ispir, M., Jain, V., Koc, L., Koo, C. Y., Lew, L., Mewald, C., Modi, A. N., Polyzotis, N., Ramesh, S., Roy, S., Whang, S. E., Wicke, M., … Zinkevich, M. (2017). TFX: A TensorFlow-based production-scale machine learning platform. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1387–1395). ACM. DOI: 10.1145/3097983.3098021. https://storage.googleapis.com/gweb-research2023-media/pubtools/4795.pdf. TFX Section 2 describes coordinated components and shared platform services. The exact metadata/artifact split here is a design choice, not proof that all orchestration systems behave this way.↩︎

  3. Sedgewick, R., & Wayne, K. (n.d.). Shortest paths: Critical path method (Author-maintained online companion to Algorithms, 4th ed., Addison-Wesley, 2011). https://algs4.cs.princeton.edu/44sp/. The authors’ algorithm companion describes finite DAG scheduling and longest paths. The guarantee concerns graph order, not availability of runtime resources or artifacts.↩︎

  4. Sedgewick, R., & Wayne, K. (n.d.). Shortest paths: Critical path method (Author-maintained online companion to Algorithms, 4th ed., Addison-Wesley, 2011). https://algs4.cs.princeton.edu/44sp/. The critical-path implementation assumes fixed-duration precedence-constrained jobs and sufficient processors. It supplies an ideal bound, not a measured ML-pipeline completion time.↩︎

  5. Apache Software Foundation. (n.d.). Best practices (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/best-practices.html. Airflow best practices recommend repeatable task outcomes and partition-specific inputs. Idempotent visible effects do not require byte-identical stochastic intermediate computation.↩︎

  6. Apache Software Foundation. (n.d.). XComs (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/xcoms.html. Airflow XComs are intended for small serializable values, with configurable backends. Passing a URI does not itself validate or retain the referenced payload.↩︎

  7. Apache Software Foundation. (n.d.). Architecture overview (Airflow 3.3.1 documentation observed September 22, 2026). https://airflow.apache.org/docs/apache-airflow/3.3.1/core-concepts/overview.html. The architecture and supervised Python Task SDK sections support the component roles and task-to-Execution-API relationship. The local in-process test path differs. No installed runtime or historical screenshot was certified against that version.↩︎

  8. Apache Software Foundation. (n.d.). Dag runs (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/dag-run.html. Dag Runs documents manual success marking and leaf-based run status, including permissive trigger-rule exceptions. The release-gate policy is an additional application requirement.↩︎

  9. Apache Software Foundation. (n.d.). Configuration reference (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/configurations-ref.html. The concurrency setting is per DAG. It does not provide shared-store coordination for external alias publication.↩︎

  10. Apache Software Foundation. (n.d.). KubernetesPodOperator (Kubernetes provider 10.22.0 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow-providers-cncf-kubernetes/stable/operators.html. KubernetesPodOperator launches pods through the Kubernetes API and does not require KubernetesExecutor. An image variable alone does not prove digest pinning, retries, or idempotent side effects.↩︎

  11. Apache Software Foundation. (n.d.). Best practices (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/best-practices.html. The official best-practices guide warns against costly or unreliable top-level code. Current process ownership is given by the separate 3.3.1 architecture citation.↩︎

  12. Apache Software Foundation. (n.d.). Tasks (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/tasks.html. The task documentation describes configured retries and exception policy. Domain classification and preserved input manifests must be supplied by application code/policy.↩︎

  13. Amazon Web Services. (n.d.). How to prevent object overwrites with conditional writes (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/conditional-writes.html. Conditional writes provide an object-level precondition. The durable publication key, winner record, read-back, and stale-attempt rules add application controls. They do not establish exactly once execution or a cross-system transaction.↩︎

  14. Kubeflow Authors. (2024, June 20). Data types (Kubeflow Pipelines documentation). https://www.kubeflow.org/docs/components/pipelines/user-guides/data-handling/data-types/. Kubeflow documents typed parameter/artifact interfaces. Type compatibility alone does not validate semantic schema, artifact quality, or upgrade compatibility.↩︎

  15. Kubeflow Authors. (2024, August 27). Create components (Kubeflow Pipelines documentation). https://www.kubeflow.org/docs/components/pipelines/user-guides/components/. Kubeflow component definitions package logic and I/O for reuse. Committed publication and validation checks are additional platform requirements.↩︎

  16. Moreau, L., & Missier, P. (Eds.). (2013, April 30). PROV-DM: The PROV data model. W3C Recommendation. https://www.w3.org/TR/2013/REC-prov-dm-20130430/. PROV-DM defines use, generation, derivation, and responsible agents. It does not prescribe this five-system split or prevent partial cross-system writes.↩︎

  17. Kubernetes Authors. (2025, May 22). Multi-tenancy (Kubernetes documentation). https://kubernetes.io/docs/concepts/security/multi-tenancy/. Kubernetes documents RBAC, network/storage controls, quotas, and trust-sensitive isolation. NetworkPolicy requires an enforcing plugin, and namespace membership alone is not a hostile-code boundary.↩︎

  18. Amazon Web Services. (n.d.). Examples of Amazon S3 bucket policies (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/example-bucket-policies.html. S3 bucket policies illustrate principal/action/resource authorization. Storage permissions and trace-store authorization are separate from scheduler capacity quotas.↩︎