References
Lists the external research and dated official documentation cited by the guide.
This section lists each external source cited in the book once. Chapter footnotes connect the sources to specific statements and explain the important limits of their evidence.
- Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems 30. https://arxiv.org/abs/1706.03762.
- Brown, T. B., Mann, B., Ryder, N., Subbiah, M., Kaplan, J., Dhariwal, P., Neelakantan, A., Shyam, P., Sastry, G., Askell, A., Agarwal, S., Herbert-Voss, A., Krueger, G., Henighan, T., Child, R., Ramesh, A., Ziegler, D. M., Wu, J., Winter, C., . . . Amodei, D. (2020). Language models are few-shot learners. In Advances in Neural Information Processing Systems 33 (pp. 1877–1901). https://arxiv.org/abs/2005.14165.
- National Institute of Standards and Technology. (2024, July). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.600-1.
- Sennrich, R., Haddow, B., & Birch, A. (2016). Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (pp. 1715–1725). Association for Computational Linguistics. https://arxiv.org/abs/1508.07909.
- Hugging Face. (n.d.). Components (Tokenizers documentation). Retrieved September 21, 2026, from https://huggingface.co/docs/tokenizers/components.
- PyTorch. (n.d.). Softmax (PyTorch 2.14 documentation). Retrieved September 16, 2026, from https://docs.pytorch.org/docs/2.14/generated/torch.nn.Softmax.html.
- Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., & Lowe, R. (2022). Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems 35 (pp. 27730–27744). https://arxiv.org/abs/2203.02155.
- Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct preference optimization: Your language model is secretly a reward model. In Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2305.18290.
- Hugging Face. (n.d.). Chat templates (Transformers v5.17.0 documentation). Retrieved September 19, 2026, from https://huggingface.co/docs/transformers/v5.17.0/en/chat_templating.
- OpenAI. (n.d.). Reasoning models. Retrieved September 19, 2026, from https://developers.openai.com/api/docs/guides/reasoning.
- Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E. H., Le, Q. V., & Zhou, D. (2022). Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems 35. https://arxiv.org/abs/2201.11903.
- Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2305.04388.
- OpenAI. (n.d.). Streaming API responses. Retrieved September 19, 2026, from https://developers.openai.com/api/docs/guides/streaming-responses.
- OpenAI. (n.d.). Mid-turn steering. Retrieved September 19, 2026, from https://developers.openai.com/api/docs/guides/steering.
- Bommasani, R., Liang, P., & Lee, T. (2023). Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525(1), 140–146. https://doi.org/10.1111/nyas.15007.
- Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (pp. 79–90). ACM. https://doi.org/10.1145/3605764.3623985.
- Ribeiro, M. T., Wu, T., Guestrin, C., & Singh, S. (2020). Beyond accuracy: Behavioral testing of NLP models with CheckList. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 4902–4912). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.442.
- Lu, Y., Bartolo, M., Moore, A., Riedel, S., & Stenetorp, P. (2022). Fantastically ordered prompts and where to find them: Overcoming few-shot prompt order sensitivity. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 8086–8098). Association for Computational Linguistics. https://aclanthology.org/2022.acl-long.556/.
- JSON Schema. (2020). JSON Schema core and validation specifications (Draft 2020-12). https://json-schema.org/specification.
- Pydantic. (n.d.). Models: Pydantic validation documentation. Retrieved September 16, 2026, from https://pydantic.dev/docs/validation/latest/concepts/models/.
- Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). Lost in the middle: How language models use long contexts. Transactions of the Association for Computational Linguistics, 12, 157–173. https://doi.org/10.1162/tacl_a_00638.
- Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., Zhang, Y., & Ginsburg, B. (2024). RULER: What’s the real context size of your long-context language models? arXiv. https://arxiv.org/abs/2404.06654.
- Tabassi, E. (2023, January). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). National Institute of Standards and Technology. https://doi.org/10.6028/NIST.AI.100-1.
- Pineau, J., Vincent-Lamarre, P., Sinha, K., Larivière, V., Beygelzimer, A., d’Alché-Buc, F., Fox, E., & Larochelle, H. (2021). Improving reproducibility in machine learning research: A report from the NeurIPS 2019 reproducibility program. Journal of Machine Learning Research, 22(164), 1–20. https://jmlr.org/papers/v22/20-303.html.
- Efron, B. (1979). Bootstrap methods: Another look at the jackknife. The Annals of Statistics, 7(1), 1–26. https://doi.org/10.1214/aos/1176344552.
- Bouthillier, X., Delaunay, P., Bronzi, M., Trofimov, A., Nichyporuk, B., Szeto, J., Sepah, N., Raff, E., Madan, K., Voleti, V., Kahou, S. E., Michalski, V., Serdyuk, D., Arbel, T., Pal, C., Varoquaux, G., & Vincent, P. (2021). Accounting for variance in machine learning benchmarks. arXiv. https://arxiv.org/abs/2103.03098.
- Papineni, K., Roukos, S., Ward, T., & Zhu, W.-J. (2002). BLEU: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics (pp. 311–318). Association for Computational Linguistics. https://aclanthology.org/P02-1040/.
- Lin, C.-Y. (2004). ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out (pp. 74–81). Association for Computational Linguistics. https://aclanthology.org/W04-1013/.
- Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. de O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., Ray, A., Puri, R., Krueger, G., Petrov, M., Khlaaf, H., Sastry, G., Mishkin, P., Chan, B., Gray, S., . . . Zaremba, W. (2021). Evaluating large language models trained on code. arXiv. https://arxiv.org/abs/2107.03374.
- White, C., Dooley, S., Roberts, M., Pal, A., Feuer, B., Jain, S., Shwartz-Ziv, R., Jain, N., Saifullah, K., Dey, S., Agrawal, S., Sandha, S. S., Naidu, S., Hegde, C., LeCun, Y., Goldstein, T., Neiswanger, W., & Goldblum, M. (2024). LiveBench: A challenging, contamination-limited LLM benchmark. arXiv. https://arxiv.org/abs/2406.19314.
- Zheng, L., Chiang, W.-L., Sheng, Y., Zhuang, S., Wu, Z., Zhuang, Y., Lin, Z., Li, Z., Li, D., Xing, E. P., Zhang, H., Gonzalez, J. E., & Stoica, I. (2023). Judging LLM-as-a-judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems 36 (Datasets and Benchmarks Track). https://arxiv.org/abs/2306.05685.
- Van der Lee, C., Gatt, A., van Miltenburg, E., Wubben, S., & Krahmer, E. (2019). Best practices for the human evaluation of automatically generated text. In Proceedings of the 12th International Conference on Natural Language Generation (pp. 355–368). Association for Computational Linguistics. https://aclanthology.org/W19-8643/.
- Anthropic. (2025, December 19). Introducing Bloom: An open source tool for automated behavioral evaluations. https://www.anthropic.com/research/bloom.
- Lewis, P., Perez, E., Piktus, A., Petroni, F., Karpukhin, V., Goyal, N., Küttler, H., Lewis, M., Yih, W., Rocktäschel, T., Riedel, S., & Kiela, D. (2020). Retrieval-augmented generation for knowledge-intensive NLP tasks. In Advances in Neural Information Processing Systems 33. https://arxiv.org/abs/2005.11401.
- Chen, T., Wang, H., Chen, S., Yu, W., Ma, K., Zhao, X., Zhang, H., & Yu, D. (2024). Dense X retrieval: What retrieval granularity should we use? In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing (pp. 15159–15177). Association for Computational Linguistics. https://aclanthology.org/2024.emnlp-main.845/.
- Thakur, N., Reimers, N., Rücklé, A., Srivastava, A., & Gurevych, I. (2021). BEIR: A heterogeneous benchmark for zero-shot evaluation of information retrieval models. In Advances in Neural Information Processing Systems 34. https://arxiv.org/abs/2104.08663.
- Robertson, S., & Zaragoza, H. (2009). The probabilistic relevance framework: BM25 and beyond. Foundations and Trends in Information Retrieval, 3(4), 333–389. https://doi.org/10.1561/1500000019.
- Cormack, G. V., Clarke, C. L. A., & Büttcher, S. (2009). Reciprocal rank fusion outperforms Condorcet and individual rank learning methods. In Proceedings of the 32nd International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 758–759). ACM. https://doi.org/10.1145/1571941.1572114.
- Nogueira, R., Jiang, Z., Pradeep, R., & Lin, J. (2020). Document ranking with a pretrained sequence-to-sequence model. In Findings of the Association for Computational Linguistics: EMNLP 2020 (pp. 708–718). Association for Computational Linguistics. https://aclanthology.org/2020.findings-emnlp.63/.
- Gao, T., Yen, H., Yu, J., & Chen, D. (2023). Enabling large language models to generate text with citations. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 6465–6488). Association for Computational Linguistics. https://aclanthology.org/2023.emnlp-main.398/.
- OWASP Gen AI Security Project. (2025). LLM01:2025 prompt injection. https://genai.owasp.org/llmrisk/llm01-prompt-injection/.
- Niu, C., Wu, Y., Zhu, J., Xu, S., Shum, K., Zhong, R., Song, J., & Zhang, T. (2024). RAGTruth: A hallucination corpus for developing trustworthy retrieval-augmented language models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 10862–10878). Association for Computational Linguistics. https://aclanthology.org/2024.acl-long.585/.
- Islam, P., Kannappan, A., Kiela, D., Qian, R., Scherrer, N., & Vidgen, B. (2023). FinanceBench: A new benchmark for financial question answering. arXiv. https://arxiv.org/abs/2311.11944.
- Xiao, S., Liu, Z., Zhang, P., Muennighoff, N., Lian, D., & Nie, J.-Y. (2023). C-Pack: Packaged resources to advance general Chinese embedding. arXiv. https://arxiv.org/abs/2309.07597.
- Meta AI. (2026). Faiss 1.15.0 [Software documentation]. https://github.com/facebookresearch/faiss/tree/v1.15.0.
- Saad-Falcon, J., Khattab, O., Potts, C., & Zaharia, M. (2024). ARES: An automated evaluation framework for retrieval-augmented generation systems. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1: Long Papers (pp. 338–354). Association for Computational Linguistics. https://aclanthology.org/2024.naacl-long.20/.
- Min, S., Krishna, K., Lyu, X., Lewis, M., Yih, W., Koh, P. W., Iyyer, M., Zettlemoyer, L., & Hajishirzi, H. (2023). FActScore: Fine-grained atomic evaluation of factual precision in long-form text generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. https://arxiv.org/abs/2305.14251.
- Es, S., James, J., Espinosa-Anke, L., & Schockaert, S. (2024). RAGAs: Automated evaluation of retrieval augmented generation. In Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demonstrations (pp. 150–158). Association for Computational Linguistics. https://aclanthology.org/2024.eacl-demo.16/.
- Ragas. (2025, December 10). Migrating from Ragas 0.3 to 0.4 (Ragas 0.4.1 documentation). https://docs.ragas.io/en/v0.4.1/howtos/migrations/migrate_from_v03_to_v04/.
- Amiraz, C., Cuconasu, F., Filice, S., & Karnin, Z. (2025). The distracting effect: Understanding irrelevant passages in RAG. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 18228–18258). Association for Computational Linguistics. https://aclanthology.org/2025.acl-long.892/.
- OWASP Gen AI Security Project. (2025). LLM06:2025 excessive agency. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/.
- Wu, Y., Yue, T., Zhang, S., Wang, C., & Wu, Q. (2024). StateFlow: Enhancing LLM task-solving through state-driven workflows. arXiv. https://arxiv.org/abs/2403.11322.
- Fielding, R., Nottingham, M., & Reschke, J. (2022). HTTP semantics (RFC 9110, section 9.2.2). RFC Editor. https://www.rfc-editor.org/rfc/rfc9110.html#name-idempotent-methods.
- Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2022). ReAct: Synergizing reasoning and acting in language models. arXiv. https://arxiv.org/abs/2210.03629.
- LangChain AI. (2026, June 29). LangGraph persistence documentation (commit
3edafc8c). GitHub. https://github.com/langchain-ai/docs/blob/3edafc8c187b52be4e2196cd0955dda4a7592cba/src/oss/langgraph/persistence.mdx. - Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). Tau-bench: A benchmark for tool-agent-user interaction in real-world domains. arXiv. https://arxiv.org/abs/2406.12045.
- Lu, J., Holleis, T., Zhang, Y., Aumayer, B., Nan, F., Bai, H., Ma, S., Ma, S., Li, M., Yin, G., Wang, Z., & Pang, R. (2025). ToolSandbox: A stateful, conversational, interactive evaluation benchmark for LLM tool use capabilities. In Findings of the Association for Computational Linguistics: NAACL 2025 (pp. 1160–1183). Association for Computational Linguistics. https://aclanthology.org/2025.findings-naacl.65/.
- Wang, L., Xu, W., Lan, Y., Hu, Z., Lan, Y., Lee, R. K.-W., & Lim, E.-P. (2023). Plan-and-Solve prompting: Improving zero-shot chain-of-thought reasoning by large language models. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers. Association for Computational Linguistics. https://arxiv.org/abs/2305.04091.
- Xu, B., Peng, Z., Lei, B., Mukherjee, S., Liu, Y., & Xu, D. (2023). ReWOO: Decoupling reasoning from observations for efficient augmented language models. arXiv. https://arxiv.org/abs/2305.18323.
- Shinn, N., Cassano, F., Berman, E., Gopinath, E., Narasimhan, K., & Yao, S. (2023). Reflexion: Language agents with verbal reinforcement learning. In Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2303.11366.
- Kim, S., Moon, S., Tabrizi, R., Lee, N., Mahoney, M. W., Keutzer, K., & Gholami, A. (2024). An LLM compiler for parallel function calling. In Proceedings of the 41st International Conference on Machine Learning. PMLR. https://arxiv.org/abs/2312.04511.
- Wang, X., Chen, Y., Yuan, L., Zhang, Y., Li, Y., Peng, H., & Ji, H. (2024). Executable code actions elicit better LLM agents. In Proceedings of the 41st International Conference on Machine Learning. PMLR. https://arxiv.org/abs/2402.01030.
- National Institute of Standards and Technology. (2017, September). Application container security guide (NIST Special Publication 800-190). U.S. Department of Commerce. https://doi.org/10.6028/NIST.SP.800-190.
- Prefect. (2026, July 29). FastMCP tools documentation (commit
0792ac81). GitHub. https://github.com/PrefectHQ/fastmcp/blob/0792ac812c3240a8256d44fbbc01caa97e4cb8cc/docs/servers/tools.mdx. - Model Context Protocol. (2025, November 25). Security best practices (version 2025-11-25). https://modelcontextprotocol.io/docs/2025-11-25/tutorials/security/security_best_practices.
- Model Context Protocol. (2025, November 25). Transports (version 2025-11-25). https://modelcontextprotocol.io/specification/2025-11-25/basic/transports.
- Model Context Protocol. (2025, November 25). Authorization (version 2025-11-25). https://modelcontextprotocol.io/specification/2025-11-25/basic/authorization.
- Campbell, B., Bradley, J., & Lodderstedt, T. (2020, February). Resource indicators for OAuth 2.0 (RFC 8707). RFC Editor. https://www.rfc-editor.org/rfc/rfc8707.html.
- Sakimura, N., Bradley, J., & Agarwal, N. (2015, September). Proof key for code exchange by OAuth public clients (RFC 7636). RFC Editor. https://www.rfc-editor.org/info/rfc7636/.
- Model Context Protocol. (2025, November 25). Elicitation (version 2025-11-25). https://modelcontextprotocol.io/specification/2025-11-25/client/elicitation.
- Model Context Protocol. (2026, July 28). Specification (version 2026-07-28). https://modelcontextprotocol.io/specification/2026-07-28.
- Model Context Protocol. (2026, July 28). Authorization (version 2026-07-28). https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization.
- Model Context Protocol. (2026, July 28). Roots (version 2026-07-28). https://modelcontextprotocol.io/specification/2026-07-28/client/roots.
- Trivedi, H., Balasubramanian, N., Khot, T., & Sabharwal, A. (2023). Interleaving retrieval with chain-of-thought reasoning for knowledge-intensive multi-step questions. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 10014–10037). Association for Computational Linguistics. https://aclanthology.org/2023.acl-long.557/.
- Jiang, Z., Xu, F., Gao, L., Sun, Z., Liu, Q., Dwivedi-Yu, J., Yang, Y., Callan, J., & Neubig, G. (2023). Active retrieval augmented generation. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 7969–7992). Association for Computational Linguistics. https://aclanthology.org/2023.emnlp-main.495/.
- Sumers, T. R., Yao, S., Narasimhan, K., & Griffiths, T. L. (2024). Cognitive architectures for language agents. Transactions on Machine Learning Research. https://arxiv.org/abs/2309.02427.
- Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K.-W., & Yu, D. (2025). LongMemEval: Benchmarking chat assistants on long-term interactive memory. In International Conference on Learning Representations. https://arxiv.org/abs/2410.10813.
- National Institute of Standards and Technology. (2022, February). Secure Software Development Framework (SSDF) version 1.1: Recommendations for mitigating the risk of software vulnerabilities (NIST Special Publication 800-218). U.S. Department of Commerce. https://doi.org/10.6028/NIST.SP.800-218.
- Anthropic. (2024, December 19). Building effective agents. https://www.anthropic.com/engineering/building-effective-agents.
- Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., . . . Tang, J. (2024). AgentBench: Evaluating LLMs as agents. In International Conference on Learning Representations. https://arxiv.org/abs/2308.03688.
- Nebius. (n.d.). Quickstart (Nebius Token Factory documentation). Retrieved September 23, 2026, from https://docs.tokenfactory.nebius.com/quickstart.
- Dror, R., Baumer, G., Shlomov, S., & Reichart, R. (2018). The hitchhiker’s guide to testing statistical significance in natural language processing. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Volume 1: Long Papers (pp. 1383–1392). Association for Computational Linguistics. https://doi.org/10.18653/v1/P18-1128.
- Miller, E. (2024). Adding error bars to evals: A statistical approach to language model evaluations. arXiv. https://arxiv.org/abs/2411.00640.
- Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., & Artzi, Y. (2020). BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations. https://arxiv.org/abs/1904.09675.
- Cohen, J. (1960). A coefficient of agreement for nominal scales. Educational and Psychological Measurement, 20(1), 37–46. https://doi.org/10.1177/001316446002000104.
- Malkov, Y. A., & Yashunin, D. A. (2020). Efficient and robust approximate nearest neighbor search using hierarchical navigable small world graphs. IEEE Transactions on Pattern Analysis and Machine Intelligence, 42(4), 824–836. https://doi.org/10.1109/TPAMI.2018.2889473.
- Carbonell, J., & Goldstein, J. (1998). The use of MMR, diversity-based reranking for reordering documents and producing summaries. In Proceedings of the 21st Annual International ACM SIGIR Conference on Research and Development in Information Retrieval (pp. 335–336). ACM. https://doi.org/10.1145/290941.291025.
- Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., & Clark, P. (2023). Self-Refine: Iterative refinement with self-feedback. In Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2303.17651.
- Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large language models cannot self-correct reasoning yet. In International Conference on Learning Representations. https://arxiv.org/abs/2310.01798.
- Agresti, A., & Coull, B. A. (1998). Approximate is better than “exact” for interval estimation of binomial proportions. The American Statistician, 52(2), 119–126. https://doi.org/10.1080/00031305.1998.10480550.
- Model Context Protocol. (2026, July 28). Deprecated features (version 2026-07-28). https://modelcontextprotocol.io/specification/2026-07-28/deprecated.
- Anthropic. (2025, June 13). How we built our multi-agent research system. https://www.anthropic.com/engineering/multi-agent-research-system.
- Erdogan, L. E., Lee, N., Kim, S., Moon, S., Furuta, H., Anumanchipalli, G., Keutzer, K., & Gholami, A. (2025). Plan-and-Act: Improving planning of agents for long-horizon tasks. In Proceedings of the 42nd International Conference on Machine Learning (PMLR 267, pp. 15419–15462). https://proceedings.mlr.press/v267/erdogan25a.html.
- OpenAI. (n.d.). Create chat completion (API reference). Retrieved September 24, 2026, from https://developers.openai.com/api/reference/resources/chat/subresources/completions/methods/create.
- Nebius. (n.d.). Create chat completion (Nebius Token Factory documentation). Retrieved September 24, 2026, from https://docs.tokenfactory.nebius.com/api-reference/inference/create-chat-completion.
- vLLM. (n.d.). vLLM V1 (vLLM documentation). Retrieved September 24, 2026, from https://docs.vllm.ai/en/latest/usage/v1_guide.html.
- OpenAI. (2023). GPT-4 technical report. arXiv. https://arxiv.org/abs/2303.08774.
- Tian, K., Mitchell, E., Zhou, A., Sharma, A., Rafailov, R., Yao, H., Finn, C., & Manning, C. (2023). Just ask for calibration: Strategies for eliciting calibrated confidence scores from language models fine-tuned with human feedback. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (pp. 5433–5442). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.emnlp-main.330.
- International Organization for Standardization, International Electrotechnical Commission, & Institute of Electrical and Electronics Engineers. (2018). Systems and software engineering: Life cycle processes: Requirements engineering (ISO/IEC/IEEE 29148:2018). https://www.iso.org/standard/72089.html.
- International Organization for Standardization & International Electrotechnical Commission. (2023). Software engineering: Systems and software Quality Requirements and Evaluation (SQuaRE): Quality model for AI systems (ISO/IEC 25059:2023). https://www.iso.org/standard/80655.html.
- Shankar, S., Zamfirescu-Pereira, J. D., Hartmann, B., Parameswaran, A. G., & Arawjo, I. (2024). Who validates the validators? Aligning LLM-assisted evaluation of LLM outputs with human preferences. In Proceedings of the 37th Annual ACM Symposium on User Interface Software and Technology (pp. 1–14). Association for Computing Machinery. https://doi.org/10.1145/3654777.3676450.
- Warner, A., & Davidovič, Š. (2018). Canarying releases. In B. Beyer, N. R. Murphy, D. K. Rensin, K. Kawahara, & S. Thorne (Eds.), The site reliability workbook: Practical ways to implement SRE (Chapter 16). O’Reilly Media. https://sre.google/workbook/canarying-releases/.
- Kohavi, R., Tang, D., & Xu, Y. (2020). Trustworthy online controlled experiments: A practical guide to A/B testing. Cambridge University Press. https://doi.org/10.1017/9781108653985.
- Jones, C., Wilkes, J., Murphy, N., & Smith, C. (2016). Service level objectives. In B. Beyer, C. Jones, J. Petoff, & N. R. Murphy (Eds.), Site reliability engineering: How Google runs production systems (Chapter 4). O’Reilly Media. https://sre.google/sre-book/service-level-objectives/.
- International Organization for Standardization & International Electrotechnical Commission. (2023). Information technology: Artificial intelligence: Guidance on risk management (ISO/IEC 23894:2023). https://www.iso.org/standard/77304.html.
- International Organization for Standardization & International Electrotechnical Commission. (2023). Information technology: Artificial intelligence: AI system life cycle processes (ISO/IEC 5338:2023). https://www.iso.org/standard/81118.html.
- International Organization for Standardization & International Electrotechnical Commission. (2023). Information technology: Artificial intelligence: Management system (ISO/IEC 42001:2023). https://www.iso.org/standard/81230.html.
- LangChain AI. (2026, June 29). LangGraph Graph API documentation (commit
3edafc8c). GitHub. https://github.com/langchain-ai/docs/blob/3edafc8c187b52be4e2196cd0955dda4a7592cba/src/oss/langgraph/graph-api.mdx. - ggml-org. (2026, September 12). GBNF Guide (
llama.cpp, commitacecd560). GitHub. https://github.com/ggml-org/llama.cpp/blob/acecd56032ddc34bada14a2d978f110d9c987095/grammars/README.md. Retrieved September 27, 2026. - LangChain AI. (n.d.). Splitting recursively. LangChain documentation. https://docs.langchain.com/oss/python/integrations/splitters/recursive_text_splitter. Retrieved September 27, 2026.
- Python Software Foundation. (2026). typing: Support for type hints (Python 3.14 documentation). https://docs.python.org/3.14/library/typing.html#typing.TypedDict. Retrieved September 28, 2026.
- LangChain AI. (n.d.). Graph API overview: Working with messages in graph state. https://docs.langchain.com/oss/python/langgraph/graph-api#working-with-messages-in-graph-state. Retrieved September 28, 2026.
- Model Context Protocol. (2025, November 25). Architecture (version 2025-11-25). https://modelcontextprotocol.io/specification/2025-11-25/architecture.
- Python Software Foundation. (n.d.). dataclasses: Data classes (Python 3.14 documentation). Retrieved September 28, 2026, from https://docs.python.org/3.14/library/dataclasses.html.
- Bitext. (n.d.). Bitext customer-support LLM chatbot training dataset [Dataset]. Hugging Face. Retrieved September 28, 2026, from https://huggingface.co/datasets/bitext/Bitext-customer-support-llm-chatbot-training-dataset.
- LangChain AI. (n.d.). Graph API overview. Retrieved September 28, 2026, from https://docs.langchain.com/oss/python/langgraph/graph-api.
- LangChain AI. (n.d.). Persistence. Retrieved September 28, 2026, from https://docs.langchain.com/oss/python/langgraph/persistence.
- Liang, P., Bommasani, R., Lee, T., Tsipras, D., Soylu, D., Yasunaga, M., Zhang, Y., Narayanan, D., Wu, Y., Kumar, A., Newman, B., Yuan, B., Yan, B., Zhang, C., Cosgrove, C., Manning, C. D., Ré, C., Acosta-Navas, D., Hudson, D. A., … Koreeda, Y. (2023). Holistic evaluation of language models. Transactions on Machine Learning Research. https://arxiv.org/abs/2211.09110v2. Retrieved September 28, 2026.
Publication and version dates identify the evidence reviewed for this edition. Recheck provider prices, model identifiers, library interfaces, protocol specifications, and authorization requirements before implementation.