References

The sources below support the claims identified by reader footnotes. Entries follow first reader-visible use and list each distinct source once. The footnotes retain the exact support and limits.

  1. Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W.-t., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. Advances in Neural Information Processing Systems, 33, 9459–9474. Curran Associates. https://proceedings.neurips.cc/paper/2020/hash/6b493230205f780e1bc26945df7481e5-Abstract.html.

  2. Rashkin, H., Nikolaev, V., Lamm, M., Aroyo, L., Collins, M., Das, D., Petrov, S., Tomar, G. S., Turc, I., & Reitter, D. (2023). Measuring attribution in natural language generation models. Computational Linguistics, 49(4), 777–840. DOI: 10.1162/coli_a_00486. https://doi.org/10.1162/coli_a_00486.

  3. Ji, Z., Lee, N., Frieske, R., Yu, T., Su, D., Xu, Y., Ishii, E., Bang, Y. J., Madotto, A., & Fung, P. (2023). Survey of hallucination in natural language generation. ACM Computing Surveys, 55(12), Article 248. DOI: 10.1145/3571730. https://doi.org/10.1145/3571730.

  4. Google Cloud. (2024, August 28). MLOps: Continuous delivery and automation pipelines in machine learning. Cloud Architecture Center. https://docs.cloud.google.com/architecture/mlops-continuous-delivery-and-automation-pipelines-in-machine-learning.

  5. PyTorch Contributors. (2024, November 26). Reproducibility. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/notes/randomness.html.

  6. NVIDIA. (n.d.). Installing the NVIDIA Container Toolkit. NVIDIA Container Toolkit documentation. https://docs.nvidia.com/datacenter/cloud-native/container-toolkit/latest/install-guide.html.

  7. Moreau, L., & Missier, P. (Eds.). (2013, April 30). PROV-DM: The PROV data model. W3C Recommendation. https://www.w3.org/TR/2013/REC-prov-dm-20130430/.

  8. Amazon Web Services. (n.d.). Shared responsibility model. https://aws.amazon.com/compliance/shared-responsibility-model/.

  9. Open Container Initiative. (2024). Open Container Initiative image format specification: Image configuration (v1.1.0). Linux Foundation. https://github.com/opencontainers/image-spec/blob/v1.1.0/config.md.

  10. Souppaya, M., Morello, J., & Scarfone, K. (2017, September). Application container security guide (NIST SP 800-190). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.SP.800-190. DOI: 10.6028/NIST.SP.800-190.

  11. Open Container Initiative. (2024). Open Container Initiative distribution specification (v1.1.0). Linux Foundation. https://github.com/opencontainers/distribution-spec/blob/v1.1.0/spec.md.

  12. Docker, Inc. (n.d.). Optimize cache usage in builds. Docker Docs. https://docs.docker.com/build/cache/optimize/.

  13. Docker, Inc. (n.d.). Multi-stage builds. Docker Docs. https://docs.docker.com/build/building/multi-stage/.

  14. Astral. (n.d.). Locking and syncing. uv documentation. https://docs.astral.sh/uv/concepts/projects/sync/.

  15. Docker, Inc. (n.d.). Build secrets. Docker Docs. https://docs.docker.com/build/building/secrets/.

  16. Docker, Inc. (n.d.). Multi-platform builds. Docker Docs. https://docs.docker.com/build/building/multi-platform/.

  17. vLLM project. (n.d.). Engine arguments (v0.10.2). https://docs.vllm.ai/en/v0.10.2/configuration/engine_args.html.

  18. NVIDIA. (n.d.). Environment variables. NCCL 2.24.3 documentation. https://docs.nvidia.com/deeplearning/nccl/archives/nccl_2243/user-guide/docs/env.html.

  19. Kubernetes Authors. (n.d.). Schedule GPUs. Kubernetes documentation. https://kubernetes.io/docs/tasks/manage-gpus/scheduling-gpus/.

  20. PyTorch Contributors. (2025, June 12). Automatic differentiation package: torch.autograd (PyTorch 2.8 documentation). https://docs.pytorch.org/docs/2.8/autograd.html.

  21. PyTorch Contributors. (n.d.). Optimizing Model Parameters (PyTorch Tutorials; observed 2.14.0+cu130). https://docs.pytorch.org/tutorials/beginner/basics/optimization_tutorial.html.

  22. PyTorch Contributors. (n.d.). Adam (PyTorch 2.8 documentation). https://docs.pytorch.org/docs/2.8/generated/torch.optim.Adam.html.

  23. Rajbhandari, S., Rasley, J., Ruwase, O., & He, Y. (2020). ZeRO: Memory optimizations toward training trillion parameter models (arXiv:1910.02054v3, May 13, 2020). arXiv. https://arxiv.org/abs/1910.02054v3.

  24. NVIDIA. (n.d.). Overview of NCCL. NCCL user guide. https://docs.nvidia.com/deeplearning/nccl/user-guide/docs/overview.html.

  25. NVIDIA. (n.d.). Overview. NVIDIA IMEX Service for NVLink Networks. https://docs.nvidia.com/multi-node-nvlink-systems/imex-guide/overview.html.

  26. Shoeybi, M., Patwary, M., Puri, R., LeGresley, P., Casper, J., & Catanzaro, B. (2019). Megatron-LM: Training multi-billion parameter language models using model parallelism (arXiv:1909.08053). arXiv. https://arxiv.org/abs/1909.08053.

  27. PyTorch Contributors. (n.d.). torch.utils.data. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/data.html#torch.utils.data.distributed.DistributedSampler.

  28. PyTorch Contributors. (n.d.). Distributed Data Parallel. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/notes/ddp.html.

  29. PyTorch Contributors. (n.d.). torch.distributed.fsdp.fully_shard. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/distributed.fsdp.fully_shard.html.

  30. Narayanan, D., Shoeybi, M., Casper, J., LeGresley, P., Patwary, M., Korthikanti, V. A., Vainbrand, D., Kashinkunti, P., Bernauer, J., Catanzaro, B., Phanishayee, A., & Zaharia, M. (2021). Efficient large-scale language model training on GPU clusters using Megatron-LM (arXiv:2104.04473v5). arXiv. https://arxiv.org/abs/2104.04473v5.

  31. NVIDIA. (n.d.). Parallelisms guide. Megatron Bridge documentation. https://docs.nvidia.com/nemo/megatron-bridge/latest/parallelisms.html.

  32. Fedus, W., Zoph, B., & Shazeer, N. (2022). Switch Transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120), 1–39. https://jmlr.org/papers/v23/21-0998.html.

  33. Huang, Y., Cheng, Y., Bapna, A., Firat, O., Chen, M. X., Chen, D., Lee, H., Ngiam, J., Le, Q. V., Wu, Y., & Chen, Z. (2019). GPipe: Efficient training of giant neural networks using pipeline parallelism. In Advances in Neural Information Processing Systems 32 (pp. 103–112). Curran Associates. https://proceedings.neurips.cc/paper/2019/hash/093f65e080a295f8076b1c5722a46aa2-Abstract.html.

  34. Korthikanti, V., Casper, J., Lym, S., McAfee, L., Andersch, M., Shoeybi, M., & Catanzaro, B. (2023). Reducing activation recomputation in large transformer models. In Proceedings of Machine Learning and Systems 5. https://arxiv.org/abs/2205.05198.

  35. PyTorch Contributors. (n.d.). torchrun (Elastic Launch). PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/elastic/run.html.

  36. PyTorch Contributors. (n.d.). Distributed communication package: torch.distributed. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/distributed.html.

  37. Volcano Authors. (n.d.). Gang. Volcano documentation. https://volcano.sh/docs/scheduler/plugins/gang/.

  38. Volcano Authors. (n.d.). Tutorials (Volcano documentation). https://volcano.sh/docs/gettingstarted/tutorials/.

  39. Kubernetes SIG Scheduling. (n.d.). Overview. Kueue documentation. https://kueue.sigs.k8s.io/docs/overview/.

  40. SkyPilot Team. (n.d.). SkyPilot: Manage all your AI compute. SkyPilot Docs. https://docs.skypilot.ai/en/latest/docs/index.html.

  41. SchedMD. (n.d.). Slurm workload manager: Overview. https://slurm.schedmd.com/overview.html.

  42. Nebius. (n.d.). Soperator: Run Slurm in Kubernetes https://github.com/nebius/soperator.

  43. Swanson, M., Bowen, P., Phillips, A. W., Gallup, D., & Lynes, D. (2010, May; updated November 11, 2010). Contingency planning guide for federal information systems (NIST SP 800-34 Rev. 1). National Institute of Standards and Technology. https://nvlpubs.nist.gov/nistpubs/Legacy/SP/nistspecialpublication800-34r1.pdf.

  44. PyTorch Contributors. (2025, June 16). Distributed Checkpoint: torch.distributed.checkpoint. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/distributed.checkpoint.html.

  45. Chowdhery, A., Narang, S., Devlin, J., Bosma, M., Mishra, G., Roberts, A., Barham, P., Chung, H. W., Sutton, C., Gehrmann, S., Schuh, P., Shi, K., Tsvyashchenko, S., Maynez, J., Rao, A., Barnes, P., Tay, Y., Shazeer, N., Prabhakaran, V., … Fiedel, N. (2023). PaLM: Scaling language modeling with pathways. Journal of Machine Learning Research, 24(240), 1–113. https://jmlr.org/papers/v24/22-1144.html.

  46. NVIDIA. (n.d.). NVIDIA System Management Interface. https://docs.nvidia.com/deploy/nvidia-smi/.

  47. Axboe, J., & fio contributors. (n.d.). fio: Flexible I/O tester (rev. 3.42, documentation build 3.42-115-gcd29). https://fio.readthedocs.io/en/latest/fio_doc.html.

  48. Amazon Web Services. (n.d.). Understanding archive retrieval options (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/restoring-objects-retrieval-options.html.

  49. Amazon Web Services. (n.d.). What is Amazon S3? Amazon Simple Storage Service User Guide. https://docs.aws.amazon.com/AmazonS3/latest/userguide/Welcome.html#ConsistencyModel.

  50. Amazon Web Services. (n.d.). RenameObject (Amazon S3 API 2006-03-01). https://docs.aws.amazon.com/AmazonS3/latest/API/API_RenameObject.html.

  51. Amazon Web Services. (n.d.). Appending data to objects in directory buckets (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/directory-buckets-objects-append.html.

  52. PyTorch Contributors. (n.d.). A guide on good usage of non_blocking and pin_memory() in PyTorch (PyTorch Tutorials, observed 2.14.0+cu130). https://docs.pytorch.org/tutorials/intermediate/pinmem_nonblock.html.

  53. Amazon Web Services. (n.d.). How to prevent object overwrites with conditional writes (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/conditional-writes.html.

  54. Amazon Web Services. (n.d.). Enforce conditional writes on Amazon S3 buckets (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/conditional-writes-enforce.html.

  55. Amazon Web Services. (n.d.). Retrieving object versions from a versioning-enabled bucket (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/RetrievingObjectVersions.html.

  56. Hugging Face. (n.d.). Tokenizer (Transformers main documentation). https://huggingface.co/docs/transformers/main/en/main_classes/tokenizer.

  57. Hugging Face. (n.d.). Text generation (Transformers main documentation). https://huggingface.co/docs/transformers/main/en/llm_tutorial.

  58. Jones, C., Wilkes, J., & Murphy, N. (with Smith, C.). (2016). Service level objectives. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems. O’Reilly Media. https://sre.google/sre-book/service-level-objectives/.

  59. Zhong, Y., Liu, S., Chen, J., Hu, J., Zhu, Y., Liu, X., Jin, X., & Zhang, H. (2024). DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation (pp. 193–210). USENIX Association. https://www.usenix.org/conference/osdi24/presentation/zhong-yinmin.

  60. Yu, G.-I., Jeong, J. S., Kim, G.-W., Kim, S., & Chun, B.-G. (2022). Orca: A distributed serving system for transformer-based generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation (pp. 521–538). USENIX Association. https://www.usenix.org/conference/osdi22/presentation/yu.

  61. vLLM project. (2025). Optimization and tuning (v0.10.2; V1 guidance). https://docs.vllm.ai/en/v0.10.2/configuration/optimization.html.

  62. Kwon, W., Li, Z., Zhuang, S., Sheng, Y., Zheng, L., Yu, C. H., Gonzalez, J. E., Zhang, H., & Stoica, I. (2023). Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (pp. 611–626). ACM. DOI: 10.1145/3600006.3613165. https://arxiv.org/abs/2309.06180.

  63. Zheng, L., Yin, L., Xie, Z., Sun, C., Huang, J., Yu, C. H., Cao, S., Kozyrakis, C., Stoica, I., Gonzalez, J. E., Barrett, C., & Sheng, Y. (2024). SGLang: Efficient execution of structured language model programs (arXiv:2312.07104v2). arXiv. Author manuscript.

  64. Ainslie, J., Lee-Thorp, J., de Jong, M., Zemlyanskiy, Y., Lebron, F., & Sanghai, S. (2023). GQA: Training generalized multi-query transformer models from multi-head checkpoints. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 4895–4901). Association for Computational Linguistics. DOI: 10.18653/v1/2023.emnlp-main.298. https://aclanthology.org/2023.emnlp-main.298/.

  65. Frantar, E., Ashkboos, S., Hoefler, T., & Alistarh, D. (2023). GPTQ: Accurate post-training quantization for generative pre-trained transformers. In The Eleventh International Conference on Learning Representations (ICLR 2023). https://arxiv.org/abs/2210.17323.

  66. Lin, J., Tang, J., Tang, H., Yang, S., Chen, W.-M., Wang, W.-C., Xiao, G., Dang, X., Gan, C., & Han, S. (2024). AWQ: Activation-aware weight quantization for LLM compression and acceleration (arXiv:2306.00978). Published in Proceedings of Machine Learning and Systems 6. https://arxiv.org/abs/2306.00978.

  67. Leviathan, Y., Kalman, M., & Matias, Y. (2023). Fast inference from transformers via speculative decoding. In Proceedings of the 40th International Conference on Machine Learning (PMLR 202, pp. 19274–19286). https://arxiv.org/abs/2211.17192.

  68. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2022). LoRA: Low-rank adaptation of large language models. In The Tenth International Conference on Learning Representations (ICLR 2022). https://arxiv.org/abs/2106.09685.

  69. SGLang project. (n.d.). SGLang model gateway. https://docs.sglang.io/docs/advanced_features/sgl_model_gateway.

  70. Kubernetes project. (n.d.). Horizontal Pod Autoscaling. https://kubernetes.io/docs/concepts/workloads/autoscaling/horizontal-pod-autoscale/.

  71. Fu, Y., Xue, L., Huang, Y., Brabete, A.-O., Ustiugov, D., Patel, Y., & Mai, L. (2024). ServerlessLLM: Low-latency serverless inference for large language models. In 18th USENIX Symposium on Operating Systems Design and Implementation (pp. 135–153). USENIX Association. https://www.usenix.org/conference/osdi24/presentation/fu.

  72. Breck, E., Cai, S., Nielsen, E., Salib, M., & Sculley, D. (2017). The ML test score: A rubric for ML production readiness and technical debt reduction. In 2017 IEEE International Conference on Big Data (pp. 1123–1132). IEEE. DOI: 10.1109/BigData.2017.8258038. https://storage.googleapis.com/gweb-research2023-media/pubtools/4156.pdf.

  73. Karpukhin, V., Oguz, B., Min, S., Lewis, P., Wu, L., Edunov, S., Chen, D., & Yih, W.-t. (2020). Dense passage retrieval for open-domain question answering. In Proceedings of EMNLP 2020 (pp. 6769–6781). Association for Computational Linguistics. DOI: 10.18653/v1/2020.emnlp-main.550. https://aclanthology.org/2020.emnlp-main.550/.

  74. Faiss maintainers. (2024, December 3). MetricType and distances (Faiss project wiki). https://github.com/facebookresearch/faiss/wiki/MetricType-and-distances.

  75. Faiss maintainers. (2025, July 28). Faiss indexes (Faiss project wiki). https://github.com/facebookresearch/faiss/wiki/Faiss-indexes.

  76. Malkov, Yu. A., & Yashunin, D. A. (2018). Efficient and robust approximate nearest neighbor search using Hierarchical Navigable Small World graphs (arXiv:1603.09320v4). arXiv. https://arxiv.org/abs/1603.09320v4.

  77. Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. DOI: 10.1561/1500000019. https://doi.org/10.1561/1500000019.

  78. Robertson, S. E., Walker, S., Jones, S., Hancock-Beaulieu, M. M., & Gatford, M. (1995). Okapi at TREC-3. In D. K. Harman (Ed.), Overview of the Third Text REtrieval Conference (TREC-3) (NIST SP 500-225, pp. 109–126). NIST. https://www.microsoft.com/en-us/research/wp-content/uploads/2016/02/okapi_trec3.pdf.

  79. Cormack, G. V., Clarke, C. L. A., & Büttcher, S. (2009). Reciprocal rank fusion outperforms Condorcet and individual rank learning methods. In Proceedings of SIGIR 2009 (pp. 758–759). ACM. DOI: 10.1145/1571941.1572114. https://cormack.uwaterloo.ca/cormacksigir09-rrf.pdf.

  80. Nogueira, R., & Cho, K. (2020). Passage re-ranking with BERT (arXiv:1901.04085v5). arXiv. https://arxiv.org/abs/1901.04085v5.

  81. Microsoft. (2026, August 24). Security filters for trimming results in Azure AI Search (Examples use API 2026-04-01). https://learn.microsoft.com/en-us/azure/search/search-security-trimming-for-azure-search.

  82. Qdrant. (n.d.). Distributed deployment (Operations documentation). https://qdrant.tech/documentation/scaling/distributed_deployment/.

  83. Qdrant. (n.d.). Snapshots (Operations documentation). https://qdrant.tech/documentation/snapshots/.

  84. Qdrant. (n.d.). Migrate to a new embedding model with zero downtime in Qdrant (Operations documentation). https://qdrant.tech/documentation/tutorials-operations/embedding-model-migration/.

  85. Qdrant. (n.d.). Collections (Collection aliases and switching). https://qdrant.tech/documentation/manage-data/collections/.

  86. National Institute of Standards and Technology & trec_eval contributors. (n.d.). m_recall.c (Official metric implementation, main branch as checked September 22, 2026). https://github.com/usnistgov/trec_eval/blob/main/m_recall.c.

  87. Voorhees, E. M. (1999). The TREC-8 question answering track report. In E. M. Voorhees & D. K. Harman (Eds.), Proceedings of the Eighth Text REtrieval Conference (TREC-8) (NIST Special Publication 500-246, pp. 77–82). National Institute of Standards and Technology. https://trec.nist.gov/pubs/trec8/papers/qa_report.pdf.

  88. Järvelin, K., & Kekäläinen, J. (2002). Cumulated gain-based evaluation of IR techniques. ACM Transactions on Information Systems, 20(4), 422–446. DOI: 10.1145/582415.582418. https://doi.org/10.1145/582415.582418.

  89. World Wide Web Consortium. (2021, November 23). Trace context (W3C Recommendation). https://www.w3.org/TR/trace-context/.

  90. OpenTelemetry authors. (n.d.). OpenTelemetry specification: Overview. https://opentelemetry.io/docs/specs/otel/overview/.

  91. OpenTelemetry authors. (n.d.). Collector. https://opentelemetry.io/docs/collector/.

  92. OpenTelemetry authors. (n.d.). Instrumentation. https://opentelemetry.io/docs/languages/python/instrumentation/.

  93. Prometheus authors. (n.d.). Metric and label naming. https://prometheus.io/docs/practices/naming/.

  94. Widmer, G., & Kubat, M. (1996). Learning in the presence of concept drift and hidden contexts. Machine Learning, 23(1), 69–101. DOI: 10.1007/BF00116900. https://link.springer.com/article/10.1007/BF00116900.

  95. Rabanser, S., Günnemann, S., & Lipton, Z. C. (2019). Failing loudly: An empirical study of methods for detecting dataset shift (arXiv:1810.11953). arXiv. Author manuscript.

  96. OpenTelemetry authors. (2026, September 16). Semantic conventions for generative client AI spans (Development status, commit be23fcc250f72e6f740c96d05fa6f12fbee3d71a). https://github.com/open-telemetry/semantic-conventions-genai/blob/be23fcc250f72e6f740c96d05fa6f12fbee3d71a/docs/gen-ai/gen-ai-spans.md.

  97. LiteLLM project. (n.d.). CLI: Quick start. https://docs.litellm.ai/docs/proxy/quick_start.

  98. MLflow project. (n.d.). LLM tracing and agent observability. https://mlflow.org/docs/latest/genai/tracing/.

  99. MLflow project. (n.d.). ML experiment tracking. https://mlflow.org/docs/latest/ml/tracking/.

  100. DVC project. (n.d.). .dvc files. https://github.com/treeverse/dvc.org/blob/main/content/docs/user-guide/project-structure/dvc-files.md.

  101. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2024). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. DOI: 10.1162/tacl_a_00638. https://aclanthology.org/2024.tacl-1.9/.

  102. Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena (arXiv:2306.05685v4). arXiv. Author manuscript.

  103. scikit-learn developers. (n.d.). cohen_kappa_score (1.9.1 documentation). https://scikit-learn.org/stable/modules/generated/sklearn.metrics.cohen_kappa_score.html.

  104. vLLM project. (n.d.). Reproducibility (v0.10.2). https://docs.vllm.ai/en/v0.10.2/usage/reproducibility.html.

  105. Nie, Y., Williams, A., Dinan, E., Bansal, M., Weston, J., & Kiela, D. (2020). Adversarial NLI: A new benchmark for natural language understanding. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 4885–4901). Association for Computational Linguistics. DOI: 10.18653/v1/2020.acl-main.441. https://aclanthology.org/2020.acl-main.441/.

  106. Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W.-t., Koh, P., Iyyer, M., Zettlemoyer, L., & Hajishirzi, H. (2023). FActScore: Fine-grained atomic evaluation of factual precision in long form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 12076–12100). Association for Computational Linguistics. DOI: 10.18653/v1/2023.emnlp-main.741. https://aclanthology.org/2023.emnlp-main.741/.

  107. Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). Tau-bench: A benchmark for tool-agent-user interaction in real-world domains (arXiv:2406.12045v1). arXiv. https://arxiv.org/abs/2406.12045.

  108. Apache Software Foundation. (n.d.). KubernetesPodOperator (Kubernetes provider 10.22.0 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow-providers-cncf-kubernetes/stable/operators.html.

  109. Baylor, D., Breck, E., Cheng, H.-T., Fiedel, N., Foo, C. Y., Haque, Z., Haykal, S., Ispir, M., Jain, V., Koc, L., Koo, C. Y., Lew, L., Mewald, C., Modi, A. N., Polyzotis, N., Ramesh, S., Roy, S., Whang, S. E., Wicke, M., … Zinkevich, M. (2017). TFX: A TensorFlow-based production-scale machine learning platform. In Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1387–1395). ACM. DOI: 10.1145/3097983.3098021. https://storage.googleapis.com/gweb-research2023-media/pubtools/4795.pdf.

  110. Sedgewick, R., & Wayne, K. (n.d.). Shortest paths: Critical path method (Author-maintained online companion to Algorithms, 4th ed., Addison-Wesley, 2011). https://algs4.cs.princeton.edu/44sp/.

  111. Apache Software Foundation. (n.d.). Best practices (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/best-practices.html.

  112. Apache Software Foundation. (n.d.). XComs (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/xcoms.html.

  113. Apache Software Foundation. (n.d.). Architecture overview (Airflow 3.3.1 documentation observed September 22, 2026). https://airflow.apache.org/docs/apache-airflow/3.3.1/core-concepts/overview.html.

  114. Apache Software Foundation. (n.d.). Dag runs (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/dag-run.html.

  115. Apache Software Foundation. (n.d.). Configuration reference (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/configurations-ref.html.

  116. Apache Software Foundation. (n.d.). Tasks (Airflow 3.3.1 documentation observed September 16, 2026). https://airflow.apache.org/docs/apache-airflow/stable/core-concepts/tasks.html.

  117. Kubeflow Authors. (2024, June 20). Data types (Kubeflow Pipelines documentation). https://www.kubeflow.org/docs/components/pipelines/user-guides/data-handling/data-types/.

  118. Kubeflow Authors. (2024, August 27). Create components (Kubeflow Pipelines documentation). https://www.kubeflow.org/docs/components/pipelines/user-guides/components/.

  119. Kubernetes Authors. (2025, May 22). Multi-tenancy (Kubernetes documentation). https://kubernetes.io/docs/concepts/security/multi-tenancy/.

  120. Amazon Web Services. (n.d.). Examples of Amazon S3 bucket policies (Amazon S3 User Guide). https://docs.aws.amazon.com/AmazonS3/latest/userguide/example-bucket-policies.html.

  121. Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations. In Proceedings of EMNLP 2023 (pp. 6465–6488). Association for Computational Linguistics. DOI: 10.18653/v1/2023.emnlp-main.398. https://aclanthology.org/2023.emnlp-main.398/.

  122. MLflow contributors. (n.d.). MLflow model registry (MLflow 2.4.2 archived documentation). https://www.mlflow.org/docs/2.4.2/model-registry.html.

  123. Schelter, S., Lange, D., Schmidt, P., Celikel, M., Biessmann, F., & Grafberger, A. (2018). Automating large-scale data quality verification. Proceedings of the VLDB Endowment, 11(12), 1781–1794. DOI: 10.14778/3229863.3229867. https://www.vldb.org/pvldb/vol11/p1781-schelter.pdf.

  124. NVIDIA. (n.d.). Routing concepts (Dynamo 1.0.2 documentation). https://docs.nvidia.com/dynamo/v1.0.2/components/router/routing-concepts.

  125. OpenTelemetry authors. (2026, January 14). Traces (OpenTelemetry documentation). https://opentelemetry.io/docs/concepts/signals/traces/.

  126. OpenTelemetry authors. (2026, January 14). Handling sensitive data (OpenTelemetry security guidance). https://opentelemetry.io/docs/security/handling-sensitive-data/.

  127. Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection (arXiv:2302.12173v2). arXiv. Author manuscript.

  128. Sculley, D., Holt, G., Golovin, D., Davydov, E., Phillips, T., Ebner, D., Chaudhary, V., Young, M., Crespo, J.-F., & Dennison, D. (2015). Hidden technical debt in machine learning systems. Advances in Neural Information Processing Systems, 28, 2503–2511. https://papers.nips.cc/paper_files/paper/2015/file/86df7dcfd896fcaf2674f757a2463eba-Paper.pdf.