8  Controlled AI workflows

An answer supported by retrieved evidence may still leave a request unfinished. Suppose a customer asks whether an order qualifies for a refund. The application must fetch the right order, check the policy, and validate the answer. Issuing the refund would require a separate authorized operation. A workflow specifies these steps and the checks that allow the process to continue.

Eight section cards for Sections 8.1 to 8.8: a proposed transfer checked for identity, permissions, balance, limits, policy and required approval; service-enforced retry with the same idempotency key; three assumed-independent stages with 0.95 cubed approximately 0.86; one route per request; independent workers and a join; criterion feedback with revision, all-checks-pass release and terminal budget stop; approval with current-state revalidation; and alternative chains, routes and parallel work with complementary feedback and approval controls.
Figure 8.1: Chains, routes, and parallel work are alternative arrangements; feedback, validation, and policy-required approval can complement them under application control.

Chains, routing, parallel work, and feedback loops organize different kinds of task. Approval and validation can be added to any of them when the action requires those checks. These patterns can be combined without following a fixed ladder of increasing complexity.

8.1 Application-controlled workflow steps

A grounded model can answer from selected evidence but remains a probabilistic text generator. A request that needs current account data or a refund also needs database queries and service calls under application authorization. The software combines model output with tools, stored data, routing, validation, and a user interface.

The model can generate language and adapt its output to a task description, but it is not an authoritative source for current private facts, deterministic arithmetic, durable storage, or authorization. Application code and external services provide those capabilities.

  • Tool: A specific application operation with a description, typed inputs, executable code, and a result returned to the model or workflow.

  • Function calling: A model interface that returns a structured proposal for a tool name and arguments. The application validates the proposal and executes only permitted calls.

A tool call is not the side effect itself. The model proposes transfer_funds(amount=...), while the application checks identity, balance, policy, limits, and approval before executing a real transfer. Every action that exposes data or changes state needs a fresh check of the caller, resource, arguments, allowed scope, and expected effect. OWASP guidance on excessive agency1 describes why applications must limit model-driven actions and permissions. Tool results are also untrusted inputs: a web page or service response can contain instructions aimed at the model.

User request to model proposal to application checks. Approval-required cases pass through human approval; allowed-without-approval cases bypass only that decision. Both paths reach current-state revalidation. Valid cases reach service execution with an atomic check and update. Denied, rejected or invalid cases terminate at Request stopped. Neutral lines attach stages to an audit record of proposal, arguments, approval, revalidation and outcome.
Figure 8.2: The model proposes a transfer. Application checks route policy-required approval and allowed-without-approval cases through current-state revalidation before the service’s atomic check and update.

The model may contribute decisions inside each stage, but the application decides which outputs are valid and which state transitions are allowed.

The model proposes text, classifications, structured arguments, or plans. The application handles message construction, identity, permissions, routing, validation, execution, persistence, budgets, and user-visible errors. External services hold authoritative records and execute side effects through controlled interfaces.

The application records what crosses each interface. A request trace identifies which context entered the model, a tool trace identifies which validated action executed, and a state trace identifies which component changed durable data. Authorization and business rules remain in application code when the model version changes. The new model still needs tests because it may propose different actions.

8.2 Workflow state and transitions

A tool can complete one operation, but a refund request needs several operations in the right order. When the sequence of steps and branches is known, having the model choose it again for every request adds avoidable variation. A workflow encodes the reliable structure and uses model calls only where language judgment is needed.

A workflow is a predefined sequence or graph of model calls, tool calls, checks, and state updates. Application code defines the available steps and the rules for moving between them. A model may supply text or a classification used by those rules. StateFlow2 is one research example of state-driven LLM workflows. Its reported gains are specific to the evaluated tasks rather than proof that one workflow design is always best.

The workflow state is the data available at one step, such as the order ID, retrieved policy, draft reply, and status. A transition runs a step, checks its result, and updates that data before selecting the next step. For the refund question, extraction produces an order ID, the account check verifies that it belongs to the caller, retrieval supplies the policy, and answer validation determines whether the draft may be returned. A failed check selects a recorded fallback instead of continuing with unchecked data.

Idempotent operation: An operation whose intended effect is the same whether an identical request is applied once or several times. RFC 91103 defines idempotency for HTTP methods. An application idempotency key identifies repeated submissions of one operation. The service uses it to recognize a retry rather than a new operation. For example, a refund service should associate one key with one set of refund parameters and one result. A retry with the same key and parameters returns the stored result instead of issuing another payment. The service must reject reuse of the key with different parameters. It must also handle simultaneous retries as one operation, using an atomic update, which checks and changes the record as one indivisible operation, or a service that provides this behavior. A separate check followed by a payment can still duplicate the effect if two requests pass the check together. Key retention and recovery after a partial failure must be defined by the service. Storing a key in workflow state alone does not make a payment idempotent.

A workflow can still call an LLM. Reliability comes from controlling when it is called, what inputs it receives, how outputs are checked, and which branch follows.

Plain functions can implement a short fixed chain, as in Section 8.3. A graph library becomes useful when branches, retries, and saved state make the connections hard to follow. In LangGraph, a node is a function that reads state and returns updates, and an edge specifies which node runs next.4 The application can declare Python state with TypedDict, a dataclass, or a Pydantic model. TypedDict describes intended field types but does not check values at runtime, so application checks are still required.5 A reducer specifies how to apply an update to a field. For example, a node can append one trace entry rather than replace the entire trace. A conditional edge calls an application function on the current state to choose the next node, so branch selection stays in code. Compiling the graph checks its structure and attaches a checkpointer when state must survive between calls (Section 9.6). Section 12.7 builds one complete graph.

Workflow

This method applies to tasks with known phases, stable policies, expected branches, or high-impact state changes. It is useful when the process needs clear records, time limits, and a defined response to failure.

Application code checks the required fields and conditions at each transition. The same control logic can be reused across requests, while changes to models, tools, policies, and data still need tests.

Limitation: Fixed sequential paths can perform poorly when the next useful action depends on unforeseen observations.

Workflow quality is measured at both step and outcome level. Step traces reveal parse errors, invalid transitions, retries, and branch selection, while outcome metrics show whether the complete process met quality, latency, cost, and policy targets. This makes a workflow appropriate for repeatable processes whose control logic benefits from software-managed execution.

8.3 Workflow control patterns

A workflow supplies a controlled sequence. One large model call may mix extraction, reasoning, transformation, and validation so thoroughly that failures cannot be located. Prompt chaining is a workflow pattern that splits a task into steps whose intermediate outputs can be typed, inspected, and retried. Routing, parallel branches, request-dependent task splitting, and feedback loops solve different control problems in the following sections. Anthropic’s practitioner guide6 describes these recurring patterns, while this chapter adds the control and measurement rules needed to choose among them.

In a static chain, the output of one step becomes input to the next. A customer support pipeline can extract entities, fetch account records, and synthesize an answer subject to validation. The chain is useful when the order is stable and each intermediate representation simplifies the next step.

Prompt chaining has operational trade-offs. Each step benefits from a narrower prompt and a testable output specification. However, every model call adds latency, token expense, and a potential failure point. Early errors can also propagate downstream unless validators halt or repair invalid intermediate representations immediately. Transformations that follow deterministic logic, such as string parsing, arithmetic, or schema validation, belong in code rather than a language-model call.

The additional calls also create a compounding reliability trade-off. For an illustrative three-stage chain, assume each stage succeeds independently with probability 0.95 and all three must succeed. The probability of complete success is \(0.95^3 = 0.857375\), about 0.86, before retries or validation. Shared model, prompt, or data errors can make stage failures dependent, so actual end-to-end reliability must be measured. Chaining is useful when intermediate results and schemas improve control or observability enough to justify the extra failure opportunities.

This Python fragment shows the chain’s control flow. It needs application-supplied helpers and TicketEntities, the expected fields for extracted ticket data. ticket_text must be a nonempty string, and customer_state must be a nonempty dictionary obtained from an authenticated account lookup. ALLOWED_INTENTS lists supported request types. The classifier and extractor may call models. The account validator checks required fields, customer identity, order ownership, and product type, returning an ok flag and a reason. Policy retrieval returns a list of approved passages. The final validator checks the draft against those passages and account state. The function returns a human_review result for an invalid input, unsupported intent, failed account check, empty retrieval, or invalid reply. On success it returns a reply result containing the checked draft. It sends no message and changes no customer data. Helper exceptions, including provider timeouts and malformed helper results, propagate to the caller, which must record the failure and stop rather than return a reply.

Code example: A chained workflow checks extracted customer and order details against account state before policy retrieval.

def handle_ticket(ticket_text, customer_state):
    if not isinstance(ticket_text, str) or not ticket_text.strip():
        return route_to_human(reason='empty or invalid ticket')
    if not isinstance(customer_state, dict) or not customer_state:
        return route_to_human(reason='empty or invalid account state')

    intent = classify_intent(ticket_text)
    if intent not in ALLOWED_INTENTS:
        return route_to_human(reason='unsupported intent')

    entities = extract_entities(ticket_text, schema=TicketEntities)
    validated = validate_customer_and_order(entities, customer_state)
    if not validated.ok:
        return route_to_human(reason=validated.reason)

    policy_chunks = retrieve_policy(intent, entities.product_type)
    if not policy_chunks:
        return route_to_human(reason='insufficient policy evidence')
    draft = draft_grounded_reply(ticket_text, policy_chunks)
    checked = validate_reply(draft, policy_chunks, customer_state)
    if not checked.ok:
        return route_to_human(reason=checked.reason)
    return {'status': 'reply', 'reply': draft}

classify_intent and extract_entities may use models, while validate_customer_and_order and validate_reply enforce application rules. Each failed check returns through route_to_human, which produces a human_review status and reason. For an illustrative ticket about returning order 42, the trace is classification, extraction, account validation, policy retrieval, drafting, and reply validation. If order 42 belongs to another customer, execution ends at the account check and no policy retrieval or drafting occurs.

8.4 Specialized routing

Prompt chaining is effective when every request follows the same ordered steps. A mixed workload may contain FAQs, calculations, document questions, and out-of-scope requests that need different tools and policies. Routing classifies a request into a small explicit set and sends it to a specialized workflow.

A router can be deterministic for known patterns or model-based for semantic categories. Its output must be one of the route labels supported by the application, or a fallback status. If a model supplies a confidence score, labeled cases must establish whether that score predicts routing accuracy before it controls escalation. Each route label identifies a worker with defined permissions and an input contract. Routing to a worker never grants permission to perform its actions.

An incoming request goes to a router labelled Selects one branch in this example. Four possible branches lead to FAQ with an insufficient-evidence response, account operation with required approval, data analysis with scope clarification, or out of scope with supported capabilities. Each lower field is headed Check / response, so required approval is not presented as fallback.
Figure 8.3: In this example, the router selects one request branch. Account operations retain required approval; the other route cards show evidence, clarification, or decline responses.

In this example, each request takes one branch. A request containing several intents needs clarification or a separate decomposition step rather than an arbitrary single label. An unsupported or ambiguous request takes a decline, clarification, or review path.

Table 8.1: A route is useful when each label leads to a concrete worker and error recovery path.
Route Input format Worker Required checks or fallback
FAQ Known policy question RAG answer workflow Insufficient-evidence response
Account operation Authenticated user and structured request Service workflow with required approval Deny unauthorized actions. Review unresolved cases
Data analysis Authorized dataset question Query planner and typed tools Scope clarification
Out of scope Unsupported domain or request Decline path Explain supported capabilities

Router evaluation uses labeled requests, per-class precision and recall, a confusion matrix, and downstream cost. A semantically plausible misroute can still be harmful when it enters a branch with different permissions or evidence sources. Confidence thresholds and deterministic prechecks can reserve ambiguous cases for clarification or human review.

8.5 Parallel work and result joining

Routing chooses one path when request types are mutually exclusive. Some tasks contain independent subtasks whose results are all needed, so choosing only one branch would be incomplete. Parallelization runs independent work concurrently. An orchestrator-worker pattern decomposes a request, assigns the subtasks to workers, and combines their results.

Fixed parallel branches are declared in advance. An orchestrator chooses the subtasks for the current request, so the two patterns differ in who determines the split. Either can use the same concurrent execution and join logic.

Parallelization is appropriate when subtasks do not depend on one another’s intermediate state. Examples include concurrent document translation, simultaneous safety checks, and multi-corpus retrieval. A join is the step that waits for required branch results and applies a rule for combining them. For example, independently translated document sections can be placed back in their original order. A safety check that needs the finished translation must run after translation, not as an independent branch. In the illustrated retrieval example, one question needs evidence from three authorized corpora. The orchestrator checks the subqueries and assignments, the workers search independently, and each returns passages with source IDs and status for the join.

One question reaches an orchestrator that checks subqueries and assigns corpora. Three independent workers search assigned corpora and each returns passages, IDs and status. Their arrows join a step that checks required results and conflicts before combining valid evidence. An orange missing-or-failed-at-deadline branch ends at Incomplete evidence / review.
Figure 8.4: An orchestrator checks subqueries for one question, assigns independent corpus searches, and joins valid evidence; a missing or failed required result at the deadline stops synthesis for incomplete-evidence handling or review.

Each worker receives the context needed for its subtask and an output specification. A narrow prompt is useful only if it retains the evidence and constraints needed to perform that work. The join step must resolve conflicts and missing results.

An orchestrator-worker pattern is useful when the way to split the work depends on the request. A model-based orchestrator adds coordination calls. Its split must be checked for missing tasks and false independence before workers run. Its extra cost is justified only when the resulting workflow improves a measured final result.

Parallel safety

Do not parallelize steps that mutate the same state without a conflict rule. Two workers can read the same old value and then overwrite one another’s updates, or issue the same payment twice. A lock, an atomic update, or separate state for each worker with a checked merge rule prevents the relevant conflict when correctly implemented.

The orchestrator defines each worker’s input, output schema, deadline, retry policy, and authorization scope. The join step handles missing, inconsistent, duplicated, and late results from each branch.

When all branches start together, parallel completion time is the slowest required branch’s time plus dispatch and joining overhead. For an illustrative three-branch task taking 2, 3, and 4 seconds, with 1 second to join, serial execution takes 10 seconds and ideal parallel execution takes 5 seconds. Shared-resource contention or limited worker slots can make the parallel run slower than that estimate. Evaluation records serial baseline time, branch durations, peak concurrency, cancellation behavior, and final synthesis quality. Parallelism is useful when saved wall time exceeds its coordination and resource cost.

8.6 Feedback and revision

Parallel work runs separate subtasks concurrently, but does not improve a draft by itself. A draft may instead need iterative improvement against a criterion that can be checked after each attempt. Evaluator-optimizer and reflection patterns add repeated feedback while preserving an explicit stop condition.

Evaluator-optimizer: A budget-limited loop that generates a candidate, grades it against an external evaluator, and revises it until the criteria pass or a budget expires. The evaluator can be deterministic, model-based, or human. It should return specific criterion failures that the next attempt can address.

Reflection: A revision pattern in which a model critiques its own or another model’s output before generating another attempt. Madaan and colleagues7 reported gains from this kind of self-feedback loop on tasks such as dialogue responses and code optimization. Without external feedback, however, Huang and colleagues8 found that self-correction on reasoning tasks often failed to help and sometimes lowered accuracy. Tests, retrieved sources, and independent graders can supply evidence the generator lacks. They still need checks for incorrect tests, irrelevant sources, and biased grading.

Evaluator-optimizer

This method applies to a candidate response and a grader that returns specific, checkable deficiencies. It is useful when quality can improve through revision and the evaluation cost is justified.

Revision is justified when the feedback identifies a checkable deficiency and repeated trials show that acting on it improves the result.

Limitation: A weak evaluator can reward revisions that raise the evaluator’s score without improving the answer.

The execution record keeps every candidate, criterion result, and stop reason. Keeping only the final answer hides whether the feedback changed the answer or the grader changed its score. Repeated trials are still needed to separate improvement from sampling variation.

Between attempts, the evaluator’s criterion-level findings become inputs to the next revision. The loop ends on a pass condition, a fixed attempt or cost budget, or a no-progress condition.

Example: One failed criterion guides one revision

Suppose a first answer says that a customer may return an item but omits the deadline stated in the retrieved policy. The evaluator checks two criteria, policy support and required conditions. It marks policy support as passing and returns missing_return_deadline for the second criterion. The next attempt receives that finding and the same policy passage, then adds the deadline with a citation. The evaluator checks the revised answer against the passage, marks both criteria as passing, and stops the loop after two attempts. The trace keeps both candidates, both criterion results, the retrieved passage, and the stop reason. This is an illustrative trace, not a measured improvement. The retained passage lets a reviewer check whether the added deadline is correct.

The evaluator becomes part of the system’s behavior and needs its own calibration. An uncalibrated judge can reward superficial verbosity or formatting that fools its scoring rule. Calibration compares the grader with held-out criterion labels and human judgments, including changed-length and changed-format answers whose correctness stays the same. Evaluation compares final quality with a one-pass baseline while also recording attempts, evaluator scores, latency, cost, and unchanged-output loops.

8.7 Validators and human approval

Feedback loops can revise low-stakes text automatically. High-consequence actions such as transferring funds or deleting records require deterministic checks, and they require human authorization when policy or the action’s risk demands it. A model score cannot supply that authorization. A human-in-the-loop gate pauses the workflow for a person’s decision. Before a high-impact action, the gate presents a proposed operation for approval, while application checks bind that approval to its parameters and relevant current state.

A human approval step should show the proposed action, parameters, evidence, expected effect, and reversible alternatives. Approval must be checked again if a change to permissions, balances, inventory, or other relevant state changes what the person approved. Revalidate identity, permissions, balances, or inventory after approval and immediately before the side effect. The service must make the final state check and the state change atomic, or reject the operation when the checked state version no longer matches. Otherwise the data can change between the check and execution.

Chapter 1 introduced safety guardrails. More generally, a guardrail is any policy or validation mechanism that restricts inputs, outputs, tool use, or state changes. It may be deterministic, model-based, or human, but permission for a high-impact action should not depend on model judgment alone.

Checks run where data enters the application, where model output is consumed, where a tool is called, and where durable state changes. Input checks validate caller credentials and payload size. Untrusted document or tool text stays separate from application instructions, but separation alone cannot guarantee that a model will ignore injected instructions. Output checks enforce the response schema and allowed values and check for sensitive data before release. Tool checks enforce permissions and argument limits. Tools that execute code also need a restricted execution environment. Before a side effect, state checks compare the current data with the approved proposal.

Approval is scoped

User consent for one visible action does not authorize a different amount, recipient, file, or future action. Store the approved parameters and reject execution when they no longer match.

Human approval is a time-limited authorization event, not a general transfer of control. The approval record stores the identity and state against which approval was granted, so execution can revalidate them immediately before the side effect.

Input checks run before the model request, output checks after generation, and tool and state checks before execution. After approval, the application repeats the checks that depend on current identity, permissions, arguments, and state. The audit record links proposal, approval, revalidation, execution, and outcome so a later review can distinguish an invalid proposal from an expired approval or failed tool.

8.8 The smallest sufficient workflow

The workflow patterns introduced in this chapter include chains, routes, parallel work, orchestrators, evaluation loops, approval, and validation. Each added workflow pattern creates more calls, branches, coordination, or failure paths. Chains, routes, and parallel branches are alternatives for arranging work. Feedback, approval, and validation can be added to those arrangements. Selection depends on which steps need earlier results, how quality can be checked, what actions change external state, and measured time or cost.

The conditions in Table 8.2 identify when the starting design needs review. They do not by themselves justify a more autonomous agent. A change must be compared with the simpler baseline on task quality, recovery, latency, and cost. Required permissions, validation, and approval apply even when they add time or cost. A simpler workflow remains the baseline.

Table 8.2: Workflow choice begins with the task dependency and risk structure rather than the desired architecture label.
Task property Preferred starting pattern Escalation condition
Known ordered steps Prompt chain or deterministic workflow Runtime observations must change the next action
Distinct request classes Router plus specialized workflows The route set cannot represent the task safely
Independent subtasks Parallel branches and explicit join Dependencies require dynamic planning
External quality criterion Evaluator-optimizer with budget Repair or replace the evaluator when its feedback is unreliable
High-impact action Approval and deterministic gate No autonomous escalation. Increase oversight

The application records the state before and after every step, the input and output schema, model and prompt version, tool call and result, validation decision, retry, cost, and wall time. This record connects workflow debugging to controlled agent tool use in Chapter 9. Appendix C shows how workflow libraries and application code divide responsibilities.

Chapter conclusion

A workflow keeps the available steps and transition rules in application code. A controlled agent is worth testing when useful next actions depend on observations that a fixed workflow cannot handle adequately. Chapter 9 adds that decision loop while keeping permissions, execution, budgets, and stopping rules in the application.


  1. OWASP Gen AI Security Project. (2025). LLM06:2025 excessive agency. https://genai.owasp.org/llmrisk/llm062025-excessive-agency/. The guidance recommends minimum tool functionality, permissions, and autonomy, downstream authorization, and human approval for high-impact actions. It is security guidance, not a controlled comparison of these controls.↩︎

  2. Wu, Y., Yue, T., Zhang, S., Wang, C., & Wu, Q. (2024). StateFlow: Enhancing LLM task-solving through state-driven workflows. arXiv. https://arxiv.org/abs/2403.11322. The paper reports results for a state-driven workflow on its evaluated tasks. It does not prove that one state model is best for every application.↩︎

  3. Fielding, R., Nottingham, M., & Reschke, J. (2022). HTTP semantics (RFC 9110, sec. 9.2.2). RFC Editor. https://doi.org/10.17487/RFC9110. The standard defines idempotency for HTTP methods. Application idempotency keys are an added design mechanism, not part of that definition.↩︎

  4. LangChain AI. (2026, June 29). LangGraph Graph API documentation (commit 3edafc8c). GitHub. https://github.com/langchain-ai/docs/blob/3edafc8c187b52be4e2196cd0955dda4a7592cba/src/oss/langgraph/graph-api.mdx. This version documents state schemas, reducers, nodes, edges, conditional edges, and compilation. It does not compare graph libraries with plain code.↩︎

  5. Python Software Foundation. (2026). typing: Support for type hints (Python 3.14 documentation, TypedDict). https://docs.python.org/3.14/library/typing.html#typing.TypedDict. The documentation distinguishes intended dictionary field types from runtime validation.↩︎

  6. Anthropic. (2024, December 19). Building effective agents. https://www.anthropic.com/engineering/building-effective-agents. The practitioner guide describes the recurring workflow patterns used here and recommends beginning with the simplest design that meets the task. It is not a controlled comparison or a universal architecture standard.↩︎

  7. Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., & Clark, P. (2023). Self-Refine: Iterative refinement with self-feedback. In Advances in Neural Information Processing Systems 36. https://arxiv.org/abs/2303.17651. The paper reports gains from model self-feedback on its evaluated tasks. It does not show that self-critique helps without an informative criterion.↩︎

  8. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large language models cannot self-correct reasoning yet. In International Conference on Learning Representations. https://arxiv.org/abs/2310.01798. The study finds that intrinsic self-correction without external feedback often fails on its reasoning benchmarks. It does not cover corrections driven by tests, tools, or retrieved evidence.↩︎