7 Running repeatable ML workflows
Every week the policy corpus changes. An update can require document validation and new embeddings before the index is rebuilt. Evaluation then informs the release decision. If a step fails halfway or runs twice, an operator needs its recorded inputs and output state to decide how to recover. The same records let another operator continue the work safely.
Selective recovery requires a record of which work was ready, which inputs it used, and which output was committed. Dependency graphs organize tasks with explicit inputs, outputs, resources, retries, and evidence.
Chapter map: repeatable ML workflows
The sections answer four linked questions:
- 7.1–7.3: What does an orchestrator add to a script, which tasks does a dependency graph allow to run, and what must each task declare?
- 7.4–7.5: Which Airflow components parse, schedule, launch, and record a run, and how does a DAG launch validation, embedding, and evaluation in isolated pods?
- 7.6–7.7: How can a task be retried without duplicating or partially replacing results, and how are common task patterns reused?
- 7.8–7.9: Which system holds which record of a run, and how do teams share one cluster without reading each other’s data?
7.1 From manual scripts to automated pipelines
A single script can run ingestion, training, evaluation, and deployment in sequence. As the workflow grows, bespoke scripts must implement durable coordination for parallel branches, selective retries, backfills, resource placement, responsibility, and persistent run state. The workflow system supplies reusable mechanisms for that work rather than making it impossible in one process. Dependency definitions stay separate from execution, and task-attempt state persists across failures.
Pipeline: A reproducible set of tasks and dependencies that transforms versioned inputs into versioned outputs while recording execution state.
Orchestrator: A system that coordinates when and where work runs and records its execution state. Workflow orchestration also checks dependencies between tasks and records their attempts.
The cluster orchestrator in Chapter 1 places workloads on machines. A workflow orchestrator coordinates the steps of a larger job. Both can participate in one pipeline: Airflow checks whether a task is ready, then a pod operator asks Kubernetes to launch its workload.1
Orchestration is not the model code. It decides when a task is eligible, where it runs, how the system records its state, and what happens after failure. The task implementation performs domain work such as preprocessing or training. The artifact store stores large outputs. The metadata store stores orchestration state and small references.2
| Concern | Ordered script | Orchestrated pipeline |
|---|---|---|
| Failure | Restart manually or add ad-hoc checkpoints | Retry or resume the failed task under policy |
| Parallelism | Manual process management | Independent branches run concurrently |
| History | Logs scattered by process | Persistent runs, attempts, owners, and timestamps |
| Scheduling | External cron or operator | Calendars, events, backfills, and concurrency limits |
| Resources | Inherited host environment | Per-task image, CPU, memory, GPU, queue, and secrets |
| Lineage | Implicit file names | Versioned input/output references and run identity |
7.2 Scheduling task dependencies with directed acyclic graphs
The workflow separates responsibilities. The orchestrator still needs an unambiguous rule for which tasks may run and which work can happen in parallel. Tasks are nodes, and prerequisite relationships form directed edges without cycles.
- Directed Acyclic Graph (DAG): A graph whose directed edges encode dependency order and whose lack of cycles means that a topological execution order exists.3
In this data contract, A → B means B requires A’s validated output, and A does not invoke B directly. Airflow edges instead govern eligibility through task state and trigger rules, so artifact readiness needs an explicit check. Independent branches can execute concurrently. A cycle has no valid strict dependency order and is normally rejected by a DAG system. Feedback over time belongs in separate runs or external events.
\[ T_{\mathrm{DAG}} \ge \max_{p \in P} \sum_{i \in p} t_i \tag{7.1}\]
\(T_{\mathrm{DAG}}\) is the completion time of one run, \(P\) is the set of paths from a starting task to a final task, and \(t_i\) is the duration of task \(i\). The longest path, the critical path, sets the lower bound because its tasks must run one after another.
Example: scheduling duration in a branched DAG
For a finite DAG with fixed task durations and finish-before-start dependencies, the longest path is a completion-time lower bound. Equality requires enough resources and no omitted queue, transfer, or retry delays:4
- Shared start: Validation takes 4 minutes.
- Training branch: Training takes 38 minutes and model evaluation takes 8 minutes.
- Retrieval branch: Embedding and index building together take 20 minutes, and index validation takes 6 minutes.
- Shared release: The promotion gate takes 3 minutes after both branches finish.
- Path sums: Adding the task durations gives 53 minutes for the training path and 33 minutes for the retrieval path.
- Conclusion: The 53-minute training path controls the lower bound. Accelerating the 33-minute branch alone cannot reduce end-to-end time below 53 minutes.
Workflow analysis: critical path scheduling implications
Identifying workflow bottleneck dependencies directs pipeline optimization efforts:
- Doubling workers does not accelerate a strictly serial chain. With fixed dependencies and durations, shortening the longest paths reduces the lower bound. Tied critical paths may need joint improvement, and changed resource waits or task internals require a new calculation.
- A slow off-path task can still consume capacity and delay later runs, so capacity and critical-path analysis remain separate.
7.3 Defining immutable task inputs and outputs
The DAG states dependency order. A box named ‘train’ is still too vague to retry safely or reproduce after an image, dataset, or configuration changes. Every task has immutable inputs, explicit outputs, resource needs, failure behavior, and validation.
Task specification: The complete contract for one pipeline task: identity, immutable inputs, outputs, runtime, resources, retry behavior, validation, and side effects.
Idempotent task: A task whose repeated execution with the same declared inputs reaches the same externally visible result without corrupting state or duplicating effects.5
Declared task fields let both the scheduler and the next task check which inputs and outputs belong to this run.
| Task field | Example | Why it matters |
|---|---|---|
| Identity | train_model:v7 plus code commit |
Connects implementation to a run |
| Inputs | Dataset manifest, base checkpoint, config hash | Prevents dependence on mutable ‘latest’ paths |
| Outputs | Checkpoint URI, metrics URI, completion manifest | Makes downstream dependencies inspectable |
| Runtime | Immutable image and entry point | Identifies image-contained dependencies; record host and external dependencies separately |
| Resources | 8 GPUs, 96 CPU, 1 TB memory, fabric constraint | Enables placement and capacity decisions |
| Retry policy | Two retries for temporary infrastructure failure, no retry on schema error | Separates recoverable from deterministic failure |
| Validation | Schema, counts, hashes, metric bounds | Prevents success from meaning only exit code zero |
| Side effects | Registry alias update or notification | Requires deduplication, transaction, or manual gate |
Large objects bypass the orchestration metadata database and inter-task message fields. The artifact store holds them, while tasks pass a small immutable reference plus integrity metadata.6 Each task writes to a unique attempt or generation path, validates its output, and conditionally publishes a completion pointer. Only then does it report task success. Section 7.6 applies the conditional-publication rules of Section 3.4 to task retries.
7.4 Coordinating workflow runs with Airflow
Task descriptions explain what each DAG node does. A concrete orchestrator still needs components that parse definitions, create runs, choose ready tasks, launch tasks, and expose state. Airflow’s components divide those responsibilities around the DAG file.
In Airflow 3.3.1, the separate DAG processor parses definitions, the scheduler contains the executor, and a triggerer is needed when deferrable operators are used. The API server supports the user interface and REST API as well as the Execution API used to report task state. In a supervised Python Task SDK run, the supervisor communicates with that Execution API on the task’s behalf. Task code does not directly access the metadata database. Earlier Airflow layouts can differ from these documented roles.7
| Airflow component | Ownership |
|---|---|
| DAG processor | Imports Python definitions and serializes valid DAG structure |
| Scheduler | Creates scheduled runs, evaluates dependencies, and queues eligible task instances |
| Executor | Within the scheduler, dispatches tasks to the selected local or remote execution backend |
| Workers / task pods | Task execution in the selected environment with reported state |
| Metadata database | Persists run, task, attempt, and scheduling state; secrets can use external backends |
| API server | Serves the UI and REST API, and the Execution API used by supervised task runtimes to communicate state |
| Artifact/log backends | Task logs and large domain outputs outside scheduler state |
| Triggerer (when used) | Runs triggers for supported deferrable operators |
The UI is a view of persisted state, not the execution engine. A success state can follow normal operator completion or manual marking, so it does not prove that domain checks ran or that the model is good. DAG-run status also depends on leaf states and trigger rules: a permissive successful leaf can mask an earlier failure. Release gates therefore consume explicit validation evidence, not merely a green run.8 Domain validation belongs inside task descriptions and release gates.
7.5 Running complete ML workflows in Airflow
The control-plane roles and task description are explicit. The Python definition describes dependencies without performing expensive work during import. Lightweight DAG parsing leaves domain code to isolated task runtimes.
Code example: Airflow dependency/runtime skeleton with external constants. No explicit retries or complete publication protocol are shown; KPO does not require KubernetesExecutor.
from airflow import DAG
from airflow.providers.cncf.kubernetes.operators.pod import (
KubernetesPodOperator,
)
from datetime import datetime
with DAG(
dag_id="versioned_rag_release",
start_date=datetime(2026, 1, 1),
schedule=None,
catchup=False,
max_active_runs=1,
tags=["mlops", "rag"],
) as dag:
validate = KubernetesPodOperator(
task_id="validate_sources",
image=PIPELINE_IMAGE,
cmds=["python", "-m", "pipeline.validate"],
arguments=["--manifest", DATASET_MANIFEST],
)
embed = KubernetesPodOperator(
task_id="build_embeddings",
image=PIPELINE_IMAGE,
cmds=["python", "-m", "pipeline.embed"],
arguments=[
"--manifest", DATASET_MANIFEST,
"--embedding-version", EMBEDDING_VERSION,
],
)
evaluate = KubernetesPodOperator(
task_id="evaluate_retrieval",
image=PIPELINE_IMAGE,
cmds=["python", "-m", "pipeline.evaluate"],
arguments=["--suite", EVAL_SUITE, "--candidate", COLLECTION_ID],
)
validate >> embed >> evaluateCode walkthrough: Airflow dependency and runtime skeleton
This skeleton defines task order and isolated runtimes. Later task contracts provide the output checks, retry rules, and publication steps needed for safe recovery:
- DAG definition: The
DAGobject stores scheduling and concurrency policy.schedule=Nonedisables automatic time scheduling, while runs can be triggered through the UI, API, command-line tool, or other control paths. - Concurrency guard:
max_active_runs=1limits overlapping runs of this DAG. It is not a lock against other DAGs, operators, or delayed external actions changing the production alias.9 - Task runtime: Each
KubernetesPodOperatorlaunches a pod for domain work after DAG import and does not require KubernetesExecutor.PIPELINE_IMAGEmust actually be digest-pinned, and host/runtime dependencies remain separate.10 - Artifact lineage: Task arguments carry dataset, embedding, evaluation-suite, and candidate-collection identities.
- Dependency state: The shift operators create validation → embedding → evaluation edges. Airflow records each task instance and attempt.
- Result: The scheduler owns eligibility and task state, while isolated containers own validation, embedding, and evaluation work.
- Limits: The uppercase constants and domain programs are supplied externally. This snippet has no explicit retry policy. A production DAG still needs explicit output manifests, retries by failure class, secrets and service accounts, resource requests, timeouts, alerts, and a separate gated promotion task.
Orchestration principle: DAG import performance safeguards
Maintaining lightweight DAG definitions protects orchestrator scheduler performance:
- Import-time dataset downloads, GPU initialization, external API calls, and large namespace scans slow parsing and can make scheduler health depend on external systems.
- The DAG processor repeatedly parses definitions. Expensive import-time work can delay parsing and scheduling visibility, while one failed import does not imply that every DAG is broken.11
7.6 Safe re-execution and idempotent tasks
The example DAG can launch isolated tasks. Retries and historical runs will repeat code, so ambiguous partitions and irreversible side effects can duplicate or overwrite production state. Each run is bound to a logical data interval or immutable release identity, while computation and promotion remain separate.
| Mechanism | Safe pattern | Failure to avoid |
|---|---|---|
| Retry | Attempt-scoped output, validation, then one committed generation | Appending duplicate rows or partially replacing a model |
| Backfill | The run’s logical interval and versioned upstream snapshot | Reading today’s mutable input for an old date |
| Catchup | Enable only when every historical interval is meaningful and resourced | Accidental flood of years of runs |
| Sensor / wait | Supported deferrable operator plus triggerer and timeout; pool policy remains configurable | Occupying a worker while polling indefinitely |
| External API | Idempotency key with persisted response identity | Duplicate billing, notification, or deployment |
| Promotion | Policy gate and conditional shared-store update after validation | A retried compute task silently deploys |
Retry policies distinguish temporary infrastructure failures from code or data failures that will recur with the same inputs. Exponential backoff multiplies the delay between successive attempts, usually up to a configured limit. This delay gives a temporarily unavailable dependency time to recover. It does not fix a schema mismatch. After the retry budget, the task state, relevant logs, input identities, and responsible component remain visible without reconstructing the run from a terminal.
Idempotency walkthrough: safe partition task retry without duplicate writes
The protocol uses staged immutable output and conditional publication, not a formal two-phase commit protocol:
- Attempt 1: Partition 07 writes to an attempt-scoped path, loses its worker, and remains uncommitted.
- Eligibility: The configured retry policy or task code classifies the failure as retryable and preserves the same logical interval and input manifest. The scheduler does not infer arbitrary domain failure classes.12
- Attempt 2: The task writes a new generation, validates row counts and hashes, and emits one completion record keyed by task, interval, and input identity.
- Commit: A coordinator applies the conditional-publication rule of Section 3.4 with a key made of task, interval, and input identity. A conditional create selects one immutable winner, a lost response is resolved by reading the winner record back before any retry, and a late attempt cannot overwrite the winner.13
- Conclusion: Attempt-scoped outputs and one publication step keep retries from duplicating or partially replacing state.
7.7 Building reusable task templates
Retry behavior and artifact publication are now safe. Teams still lose consistency when every DAG reimplements data validation, distributed training, inference, evaluation, and registration differently. Parameterized task blocks package common responsibilities through stable inputs, outputs, checks, and platform defaults.14
For example, a reusable validation block accepts an input-manifest ID, a digest-pinned image, a resource request, and a retry class. It validates the input, writes an output manifest, records telemetry, and passes its publication check before a downstream task can use the result. The task-specific part is the domain validation rule and its thresholds. The platform-owned part is identity, secret delivery, scheduling, retry recording, telemetry, and conditional publication.
Each reusable block has declared inputs, committed outputs, and validation checks.15
| Block | Consumes | Produces | Primary gate |
|---|---|---|---|
| Data validation | Source snapshot and schema | Validated manifest and profile | Schema, counts, drift, policy |
| Training | Dataset, base model, config, image | Checkpoint and training evidence | Finite loss, throughput, checkpoint integrity |
| Batch/online inference | Serving bundle and input set | Outputs and traces | Interface, latency, errors |
| Evaluation | Outputs, references, evaluator version | Metrics, slices, failed cases | Quality/safety thresholds |
| Registry | Validated artifact and metadata | Immutable version and lineage | Completeness; integrity and authenticity checks required by policy |
| Deployment | Approved serving bundle | Candidate endpoint and rollout state | Health, canary, rollback |
In the platform contract, a reusable block should standardize identity, evidence, secrets, resource requests, telemetry, and output publication while allowing task-specific domain parameters. This division lets a model team change architecture without reimplementing cluster scheduling or artifact lineage.
7.8 Recording artifact evidence in run manifests
The pipeline can produce versioned artifacts through reusable blocks. Five systems now record overlapping facts, and copying all data into every system obscures the authoritative record. Each evidence class has one primary system, and stable identifiers connect the records.
| System | Canonical responsibility | Cross-link |
|---|---|---|
| Orchestrator | DAG run, task attempts, dependency state, schedule, retries | Experiment run ID and artifact URIs |
| Experiment tracker | Parameters, metrics, model/evaluator versions, comparable research runs | DAG/task ID and deployment candidate |
| Artifact/registry store | Immutable files, manifests, versions, promotion aliases | Producer run and consumer release |
| Observability backend | Service logs, metrics, traces, alerts, incidents | Serving bundle, DAG run, experiment and request IDs |
| Evaluation store | Versioned cases, references, raw judgments, slice reports | Experiment, bundle, dataset and evaluator IDs |
A pipeline task can start an experiment run and emit its ID. The tracker records comparable parameters and metrics. The task publishes artifact URIs and commits its status. A later deployment writes the serving-bundle ID into traces. The same identifiers let an incident move backward from one request to its bundle, evaluation report, experiment, DAG run, and source artifacts.
Architecture boundary: authority mapping across systems
The responsibility map reduces ambiguity. It does not make cross-system writes atomic:
- The orchestrator is authoritative for task attempts and dependency state, the experiment tracker is authoritative for comparable parameters and metrics, the artifact registry is authoritative for immutable files and promotion aliases, telemetry is authoritative for runtime observations, the evaluation store is authoritative for cases and raw judgments.
- Cross-links join these records without copying one system’s mutable state into another. A run manifest records the identifiers and the producer or consumer relationship.16
If artifact publication succeeds but a tracker update fails, a durable pending-link record keeps the missing update visible. A reconciliation task retries that link and checks the selected generation. The authoritative record and any delayed mirror remain distinguishable during recovery.
7.9 Isolating teams on shared research clusters
One pipeline has clear responsibilities and evidence. A research organization adds concurrent users, shared clusters, queues, secrets, quotas, interactive work, long-running training, and repeated evaluation. Code, compute, storage, observability, orchestration, and experiment systems collaborate while retaining distinct roles.
A shared research platform combines Git-hosted code, containerized workloads, Kubernetes scheduling, shared and object storage, telemetry, pipeline control, and experiment tracking. Interactive work and scheduled jobs can use the same immutable images and artifact identities, allowing an exploratory result to become a repeatable run.
Operators and researchers use different evidence from the same run. Pod state and task logs help operators investigate stalled or failed work. Training progress and evaluation help researchers decide whether the run is worth continuing. The same run ID connects both views.
Tenant boundary
- In this shared cluster, namespaces or projects group each tenant’s resources, while quotas and queue limits bound resource use. Each workload runs under an identity with permission to call only the required APIs. Enforced network policies limit connections, storage access policies control reads and writes, and trace-store permissions control who can inspect recorded requests. A storage prefix groups objects but does not itself enforce access. Mutually untrusted code can require stronger node, sandbox, or cluster separation.17
- Run and artifact identities remain globally traceable, while payload access follows tenant policy. A scheduler quota alone does not prevent one team from reading another team’s checkpoints or traces.18
Chapter 7 summary
Key orchestration and workflow principles established in this chapter:
- Core mechanisms: A DAG states which tasks depend on which outputs, and its critical path sets an ideal lower bound on run time under the stated DAG assumptions. Each task declares immutable inputs, outputs, resources, retries, and validation. Airflow parses definitions, schedules ready tasks, and records attempts, while domain work runs in isolated pods. Run manifests link task attempts to experiments, artifacts, and traces.
- Governing trade-offs: More workers cannot overlap tasks whose dependencies require them to run in sequence. They can reduce resource waits or shorten a task that can use more workers, so the critical path must be recalculated when task durations change. Reusable task blocks standardize identity, evidence, and publication at the cost of a shared platform contract. A shared cluster can improve utilization but needs separate access controls to keep tenants apart.
- Failure modes & defenses: A green run can hide a skipped check, so release gates read validation evidence. Retries write attempt-scoped outputs and publish one winner. A scheduler quota does not stop one team from reading another’s checkpoints, so storage and trace permissions are separate controls.
Chapter checkpoint
Review Questions 35–38 in Appendix B, Section B.1, to test workflow records, idempotent retries, the separation of computation from production-alias promotion, and the critical-path calculation before examining the complete operational case.
Carry-forward result
The system now expresses work as a DAG of idempotent task descriptions, stores artifacts outside control-plane metadata, joins orchestration with experiments and telemetry, and supports selective retry and recovery.
The final chapter applies every layer to the grounded-answer service and follows a constructed quality regression from detection to verified release.
Apache Software Foundation. (n.d.). KubernetesPodOperator (Kubernetes provider 10.22.0 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow-providers-cncf-kubernetes/stable/operators.html. KubernetesPodOperator launches pods through the Kubernetes API and does not require KubernetesExecutor. An image variable alone does not prove digest pinning, retries, or idempotent side effects.↩︎
Baylor, D., Breck, E., Cheng, H.-T., Fiedel, N., Foo, C. Y., Haque, Z., Haykal, S., Ispir, M., Jain, V., Koc, L., Koo, C. Y., Lew, L., Mewald, C., Modi, A. N., Polyzotis, N., Ramesh, S., Roy, S., Whang, S. E., Wicke, M., … Zinkevich, M. (2017). TFX: A TensorFlow-based production-scale machine learning platform. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1387–1395). ACM. DOI: 10.1145/3097983.3098021. https://storage.googleapis.com/gweb-research2023-media/pubtools/4795.pdf. TFX Section 2 describes coordinated components and shared platform services. The exact metadata/artifact split here is a design choice, not proof that all orchestration systems behave this way.↩︎
Sedgewick, R., & Wayne, K. (n.d.). Shortest paths: Critical path method (Author-maintained online companion to Algorithms, 4th ed., Addison-Wesley, 2011). https://algs4.cs.princeton.edu/44sp/. The authors’ algorithm companion describes finite DAG scheduling and longest paths. The guarantee concerns graph order, not availability of runtime resources or artifacts.↩︎
Sedgewick, R., & Wayne, K. (n.d.). Shortest paths: Critical path method (Author-maintained online companion to Algorithms, 4th ed., Addison-Wesley, 2011). https://algs4.cs.princeton.edu/44sp/. The critical-path implementation assumes fixed-duration precedence-constrained jobs and sufficient processors. It supplies an ideal bound, not a measured ML-pipeline completion time.↩︎
Apache Software Foundation. (n.d.). Best practices (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/best-practices.html. Airflow best practices recommend repeatable task outcomes and partition-specific inputs. Idempotent visible effects do not require byte-identical stochastic intermediate computation.↩︎
Apache Software Foundation. (n.d.). XComs (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/xcoms.html. Airflow XComs are intended for small serializable values, with configurable backends. Passing a URI does not itself validate or retain the referenced payload.↩︎
Apache Software Foundation. (n.d.). Architecture overview (Airflow 3.3.1 documentation observed September 22, 2026). https://airflow.apache.org/docs/apache-airflow/3.3.1/core-concepts/overview.html. The architecture and supervised Python Task SDK sections support the component roles and task-to-Execution-API relationship. The local in-process test path differs. No installed runtime or historical screenshot was certified against that version.↩︎
Apache Software Foundation. (n.d.). Dag runs (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/dag-run.html. Dag Runs documents manual success marking and leaf-based run status, including permissive trigger-rule exceptions. The release-gate policy is an additional application requirement.↩︎
Apache Software Foundation. (n.d.). Configuration reference (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/configurations-ref.html. The concurrency setting is per DAG. It does not provide shared-store coordination for external alias publication.↩︎
Apache Software Foundation. (n.d.). KubernetesPodOperator (Kubernetes provider 10.22.0 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow-providers-cncf-kubernetes/stable/operators.html. KubernetesPodOperator launches pods through the Kubernetes API and does not require KubernetesExecutor. An image variable alone does not prove digest pinning, retries, or idempotent side effects.↩︎
Apache Software Foundation. (n.d.). Best practices (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/best-practices.html. The official best-practices guide warns against costly or unreliable top-level code. Current process ownership is given by the separate 3.3.1 architecture citation.↩︎
Apache Software Foundation. (n.d.). Tasks (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/tasks.html. The task documentation describes configured retries and exception policy. Domain classification and preserved input manifests must be supplied by application code/policy.↩︎
Amazon Web Services. (n.d.). How to prevent object overwrites with conditional writes (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/conditional-writes.html. Conditional writes provide an object-level precondition. The durable publication key, winner record, read-back, and stale-attempt rules add application controls. They do not establish exactly once execution or a cross-system transaction.↩︎
Kubeflow Authors. (2024, June 20). Data types (Kubeflow Pipelines documentation). https://www.kubeflow.org/docs/components/pipelines/user-guides/data-handling/data-types/. Kubeflow documents typed parameter/artifact interfaces. Type compatibility alone does not validate semantic schema, artifact quality, or upgrade compatibility.↩︎
Kubeflow Authors. (2024, August 27). Create components (Kubeflow Pipelines documentation). https://www.kubeflow.org/docs/components/pipelines/user-guides/components/. Kubeflow component definitions package logic and I/O for reuse. Committed publication and validation checks are additional platform requirements.↩︎
Moreau, L., & Missier, P. (Eds.). (2013, April 30). PROV-DM: The PROV data model. W3C Recommendation. https://www.w3.org/TR/2013/REC-prov-dm-20130430/. PROV-DM defines use, generation, derivation, and responsible agents. It does not prescribe this five-system split or prevent partial cross-system writes.↩︎
Kubernetes Authors. (2025, May 22). Multi-tenancy (Kubernetes documentation). https://kubernetes.io/docs/concepts/security/multi-tenancy/. Kubernetes documents RBAC, network/storage controls, quotas, and trust-sensitive isolation. NetworkPolicy requires an enforcing plugin, and namespace membership alone is not a hostile-code boundary.↩︎
Amazon Web Services. (n.d.). Examples of Amazon S3 bucket policies (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/example-bucket-policies.html. S3 bucket policies illustrate principal/action/resource authorization. Storage permissions and trace-store authorization are separate from scheduler capacity quotas.↩︎