11  Coordinating agents and managing persistent memory

One agent may struggle to select among many similar tools, keep unrelated instructions in its context, or finish independent work quickly enough. Separate agent loops can divide that work, but they also introduce transfers of control and shared state. Their value must be tested against one controlled loop, using the permissions, budgets, and outcome checks established in Chapters 8 to 10. Reusing information in later sessions adds a separate decision about what the application may store and retrieve.

Six cards match Sections 11.1 to 11.6. One-agent and divided work share quality, latency, and cost criteria without invented results. A checked handoff carries task, authority, and a return route; denial stops work. Four memory roles support selection, optional promotion, correction, and deletion. Selected working state and dated events form context. Candidate A has relevance 0.8, freshness 0.6, reliability 1.0, and score 0.78; B has 0.9, 0.2, 0.8, and score 0.67. A current delivery fact supersedes an older one, whose history remains only if permitted. Procedures pass review and tests before release, with rollback to a tested policy-compatible version.
Figure 11.1: Agent coordination and persistent memory need separate controls. The illustrative ranking favors A at 0.78 over B at 0.67 despite B having higher relevance; source reliability and freshness also contribute.

The map follows two connected questions: when to divide work among agents, and how to control information passed to another agent or reused in a later session. The memory types describe different roles for information, not a required set of four databases.

11.1 When extra agents are warranted

A support application may need billing and account-access tools with different permissions and instructions. Placing them in separate agent loops can reduce the choices each model sees. Application code must still decide which loop runs, what it may read or change, and when the whole task is complete. The extra loops add routing calls, transferred state, and duplicated context, so evaluation must test route selection, state transfer, loop detection, and incomplete results after worker failure.

  • Multi-agent system: A controlled composition of two or more agent loops with explicit communication, state access, routing, and termination rules enforced by the surrounding application.

Four recurring limitations provide concrete reasons to consider that composition:

  • Tool overload: one agent prompt can contain so many similar tools that selection degrades.
  • Context bloat: unrelated domain instructions and evidence compete in one context window.
  • Serial bottleneck: independent subtasks dominate wall-clock time when executed by one loop.
  • Separate teams or permissions: different domains need different responsible teams, permissions, release schedules, or compliance rules.

Escalation test

A decision to add agents needs a task requirement or measured single-agent limit, a prediction of which multi-agent pattern addresses it, and tests for the new failure modes.

Four topology panels. Three network peers have bidirectional links. An orchestrator sends tasks to three workers, which return results along dashed arrows. A hierarchy sends tasks through two domain coordinators to four workers, with dashed results returning upward. One swarm trace passes control from A to B to C and stops, while gray links indicate other permitted routes. A shared controls strip lists permission checks, budgets, and termination.
Figure 11.2: The patterns differ in who directs work and receives results. Solid arrows show task or control transfer, dashed arrows show returns, and gray routes show other permitted handoffs. Application checks govern permissions, budgets, and termination in every pattern.

In an orchestrator-worker pattern, a central router selects a worker and receives its result. A hierarchy adds routers for separate domains. In a network, an agent may communicate with any peer. Here, swarm means peer handoffs: the active worker selects the next worker without a central router. It does not mean that every worker runs at once. The figure sketches these possible connections, not a complete runtime graph. The table compares them using task dependencies, permissions, routing delay, and failure handling. These selection conditions are engineering guidance, not a guarantee that one pattern performs best.

Anthropic1 reported that a research system using Claude Opus 4 as lead agent and Claude Sonnet 4 workers outperformed single-agent Claude Opus 4 by 90.2 percent on its internal research evaluation. Separately, it reported about 15 times as many tokens for multi-agent systems as for chat interactions. The different baselines matter: 15 times is not the reported token ratio for the evaluated single-agent comparison. The report also notes that shared-context tasks and tasks with many dependencies between agents can fit this approach poorly.

Table 11.1: Multi-agent patterns are communication topologies, not personas.
Pattern Who selects the next agent? Conditions that can favor the pattern Risks to test
Network Each agent may select any peer Small exploratory debate or research groups A fully connected set of N agents allows N(N-1) directed peer links and many possible interactions
Orchestrator-worker A central router Central routing across specialized workers Router latency, misrouting, and supervisor drift
Hierarchical Routers at multiple domain levels Large organizations or many isolated domains Multiple routing layers, latency, and error propagation
Swarm The active worker hands off directly Stateful conversations with infrequent domain shifts Ping-pong handoffs and hidden global completion state

11.2 Bounded agent handoffs

A pattern describes the possible communication paths. A concrete system still needs a mechanism for changing the active agent loop without losing user intent, authorization, or progress. A handoff identifies the destination, transfers only the required state, records the output expected from that loop, and states where control returns.

Application code checks the destination, task, completion criteria, return route, and authorization before changing the active loop. It loads only evidence the destination may read and records the accepted handoff with the resulting state. The typed record below is an application design, not an MCP message type. MCP standardizes tool and context exchanges (Section 10.5), not this record’s task and return-control rules.

This abbreviated Python 3.10-or-later example needs application-supplied helpers and a state object. validate_destination rejects unknown workers. validate_authorization checks that the authority covers this task and destination. load_refs checks access to each reference and enforces the application’s payload-size limit. An empty evidence list is allowed for a task that needs no external evidence. Literal documents labels but does not validate them at runtime.2 The explicit checks reject empty task fields, invalid return labels, and malformed reference lists. transfer_control must save the accepted record and changed state together, or fail without applying either. Helper failures before that call leave the active agent unchanged. The fragment does not run a worker or grant permission for its later tool actions, which need the checks from Section 8.7.

Code example: The checked handoff retains the task and return route with evidence loaded under application access and size limits.

from dataclasses import dataclass
from typing import Literal

@dataclass
class Handoff:
    destination: Literal["billing", "auth", "notifier"]
    reason: str
    task: str
    completion_criteria: str
    evidence_refs: list[str]
    authorization_ref: str | None
    return_to: Literal["orchestrator", "user", "none"]

def accept_handoff(handoff: Handoff, state):
    validate_destination(handoff.destination)
    for value in (handoff.reason, handoff.task, handoff.completion_criteria):
        if not isinstance(value, str) or not value.strip():
            raise ValueError("handoff fields must be nonempty strings")
    if handoff.return_to not in {"orchestrator", "user", "none"}:
        raise ValueError("invalid return route")
    if not isinstance(handoff.evidence_refs, list) or any(
        not isinstance(ref, str) or not ref.strip()
        for ref in handoff.evidence_refs
    ):
        raise ValueError("invalid evidence references")
    validate_authorization(
        handoff.authorization_ref, handoff.task,
        destination=handoff.destination,
    )
    context = load_refs(
        handoff.evidence_refs, destination=handoff.destination,
        authorization_ref=handoff.authorization_ref,
    )
    return state.transfer_control(
        active_agent=handoff.destination,
        bounded_context=context,
        task=handoff.task,
        reason=handoff.reason,
        completion_criteria=handoff.completion_criteria,
        authorization_ref=handoff.authorization_ref,
        evidence_refs=list(handoff.evidence_refs),
        return_to=handoff.return_to,
    )

For an illustrative handoff, an orchestrator sends billing the task “check order 42’s refund eligibility”, its policy reference, and the completion criterion to return eligibility with the policy source. The accepted state retains that task, authority reference, evidence references, and the route back to the orchestrator. A rejected authority reference stops execution before evidence is loaded or control changes. After the worker returns, the orchestrator checks the criterion before marking the task complete. The table lists further failure controls, such as handoff budgets and worker-status checks, that this fragment does not implement.

Table 11.2: Handoff failures are state and control failures.
Failure Observable symptom Mitigation
Routing oscillation Agents hand control back and forth without new evidence Handoff budget, visited-route state, and loop detection
Context loss Destination repeats work or misreads user intent Typed task, evidence references, constraints, and completion criteria
Supervisor drift Central router begins solving domain work itself Narrow router outputs and separate route-accuracy evaluation
Hidden partial failure One worker fails but synthesis sounds complete Per-worker status, required-result set, and incomplete-result response
Permission amplification Worker receives broader permissions than the originating task requires Reduce permissions to the task and destination, then recheck each proposed action

A bounded handoff carries only the state needed by the next agent loop. Reusing information after the turn, session, or software version ends creates a separate storage problem: application policy must define what may persist, who may read it, how it is updated or corrected, and when it is deleted.

11.3 Memory lifecycle

The handoff state supports the next turn. Some applications must also reuse selected information in later sessions. Without explicit lifecycle rules, the application may copy every past message into durable storage and later place stale or unauthorized data into a new request.

Sumers and colleagues3 use CoALA to distinguish four kinds of memory. Their framework does not require four physical databases.

  • Working memory: The active state for the current task, including goals, plans, evidence, tool definitions, and observations. The application selects part of it for each model request.
  • Episodic memory: Timestamped records of earlier interactions, actions, and outcomes that an application may retrieve for a later request.
  • Semantic memory: Curated facts or preferences that an application stores separately from the events from which they were derived.
  • Procedural memory: Versioned prompts, tools, workflows, examples, or model adaptations that influence how the system performs a task.

Persistent records can retain information that cannot be reconstructed from the current request, such as a previously stated preference. Reuse is worthwhile when the information is relevant, permitted, and current enough for the task, and its benefit justifies storage, retrieval, and review costs.

Table 11.3: Memory types differ by representation and lifecycle.
Memory type Content Typical store Key control question
Working Active task state, messages, plan, and tool observations Application state with a selected portion in the request context What deserves scarce context tokens now?
Episodic Timestamped interactions and outcomes Event log or vector/time index Which past event is relevant and permitted?
Semantic Curated facts and preferences Profile, knowledge graph, or fact store Is the fact valid, scoped, current, and correctable?
Procedural Reusable prompts, examples, tools, workflows, and model adaptations Versioned code and instruction assets Who reviewed the procedure and which version ran?
  • Memory system: The database, index, and application logic that manages persistent state across agent sessions.
  • Capture: The process of identifying candidate information from the conversation and associating it with a source.
  • Promotion: The policy-controlled decision to store a captured candidate as a durable fact or procedure after checking its source, identity, scope, confidence, retention rule, and required human approval.

These categories help the application apply different storage and reuse rules. Separation alone enforces no rule: capture and promotion checks determine whether a past message may become a durable, reusable fact.

11.4 Working and episodic memory

A later question may need one event from a long conversation rather than the whole history. Working memory holds active task state, while episodic memory keeps dated events that the application may retrieve into that state. The request context contains only the selected portion that fits alongside instructions and reserved output space.

A larger context window does not answer the retrieval question. Irrelevant history can crowd out policy, evidence, and output space even when it technically fits.

The application selects system policy, current objectives, active plans, retrieved evidence, tool definitions, and relevant observations for the next request. It reserves output tokens separately. Summaries can reduce history size, but may lose details, so important claims should retain references to their source events.

Semantic similarity alone is weak for temporal questions. Episodic retrieval often needs user, session, time, entity, and access filters alongside embeddings.

An episodic store captures event identity, timestamps, actor roles, actions, observations, and the source record. Retrieval applies user and access filters before ranking, and time or entity filters when the question requires them. For example, “which delivery address did I give before the move?” needs the correct user’s events from before the move, not merely the most similar address message. The application loads the selected event into working memory and preserves its date in the answer evidence. A summary should retain a stable source reference while the source may be kept. If that source expires or is deleted, the summary must be removed, updated, or excluded according to the same policy.

Conversation history is user data, not free application context. The accountable product owner or data policy owner defines retention and deletion rules, and application code enforces them when storing, retrieving, expiring, and deleting records. The NIST Generative AI Profile4 supports governance of information integrity, privacy, and lifecycle risk, but it does not prescribe one universal schedule. Evaluation should measure episodic recall, irrelevant records injected into context, harm from stale facts, and compliance with the applicable rules.

11.5 Semantic memory and correction

Episodic memory preserves dated events and their outcomes. Raw events do not automatically become reliable current facts, especially when observations conflict or user preferences change. Semantic memory introduces an explicit promotion and correction process before a fact becomes reusable across sessions.

Storing extracted statements without review can create duplicate, stale, conflicting, and mis-scoped profile records. Promotion therefore needs a schema, confidence rule, scope, source record, conflict policy, and correction path.

An episode produces a candidate fact that reaches a source, consent, scope, and expiry policy check. A permitted branch reaches a semantic record directly. A review-required branch reaches human review, then either semantic storage or a not-promoted outcome. Policy denial also reaches not promoted. Supporting records, observation time, scope, expiry, and status accompany stored records. Episode retention remains conditional.
Figure 11.3: Promotion policy permits direct storage, requires human review, or rejects the candidate. Rejection creates no semantic record; the original episode remains only while its retention and deletion rules permit.

The promotion step is an application policy. The model may propose a candidate memory, but deterministic code or human review decides whether it becomes durable.

After identity, permission, status, and deletion filters remove ineligible records, the application ranks remaining records by relevance, freshness, and source reliability.

\[ S(m,q)=w_r\mathrm{relevance}(m,q)+w_f\mathrm{freshness}(m)+w_s\mathrm{reliability}(m) \tag{11.1}\]

Here, \(m\) is an eligible record and \(q\) is the request. The component functions return relevance, freshness, and reliability scores. Product weights \(w_r\), \(w_f\), and \(w_s\) balance these scores to rank candidates without overriding access rules.

Freshness needs a defined clock and a valid observation time. This example rejects future observed_at values and computes a nonnegative age in days when the record is read. With a 30-day useful-life horizon, freshness is the larger of zero and 1 - age_in_days / 30. A record observed 6 days ago receives 0.8, while one at least 30 days old receives 0 and should normally be refreshed or excluded. An expires_at rule can remove a record before ranking. The 30-day horizon is only an example. It must reflect how quickly the underlying fact changes. The application compares the ranking with reviewed retrieval choices. The score is not a probability of truth.

In the example below, each component uses a 0 to 1 scale, and the weights express the product’s chosen trade-off. Another design can use other scales or weights that do not sum to one.

Example: Rank two eligible memory records

Suppose two permitted records can help answer a question. Use weights 0.5 for relevance, 0.3 for freshness, and 0.2 for source reliability. Record A scores 0.8, 0.6, and 1.0. Its freshness of 0.6 corresponds to an age of 12 days under the 30-day example policy. Record B has higher relevance, 0.9, but lower freshness, 0.2, and reliability, 0.8.

1. Relevance contributes 0.5 × 0.8 = 0.40. 2. Freshness contributes 0.3 × 0.6 = 0.18. 3. Source reliability contributes 0.2 × 1.0 = 0.20. 4. The combined score is 0.40 + 0.18 + 0.20 = 0.78. 5. Record B scores 0.5 × 0.9 + 0.3 × 0.2 + 0.2 × 0.8 = 0.67. Result: Record A ranks above B despite B’s higher relevance.

Interpretation: The chosen weights trade some relevance for freshness and source reliability. The arithmetic does not establish that this trade-off improves answers, which needs reviewed retrieval tests. A deleted, unauthorized, or wrong-subject record remains ineligible regardless of its score.

This Python record describes a fact’s fields, not a validator. The application defines identity, category, intended-use scope, source tracking, and retention rules. source_episode_ids identifies the conversations or events from which the fact was derived. The provenance field holds references to supporting records checked during validation. The dataclass constructor requires the declared fields, but does not check their annotated types or allowed values.5 The store must validate the status, identity, scope, confidence, and dates before accepting the record.

Code example: A semantic fact record carries lifecycle metadata.

from dataclasses import dataclass
from datetime import datetime
from typing import Literal

@dataclass
class MemoryFact:
    subject_id: str
    predicate: str
    value: str
    category: str
    scope: str
    source_episode_ids: list[str]
    provenance: list[str]
    confidence: float
    observed_at: datetime
    created_at: datetime
    reviewed_at: datetime | None
    expires_at: datetime | None
    status: Literal["candidate", "active", "superseded", "deleted"]

The record describes what a fact stores. The next fragment shows when the application may store it and how deletion reaches the locations that serve it.

  • Tombstone: A small record stating that an item was deleted or replaced, without retaining the removed value. Replicas and indexes use it to avoid restoring an old copy during synchronization.

This abbreviated Python fragment uses MemoryFact and application-supplied policy, validation, store, index, cache, and audit helpers. require raises on a failed condition. validate_candidate checks the record’s values and supporting evidence. The caller then checks that validation has not changed its subject, category, or scope, and stores the whole validated record as active with the policy’s expiry. This example requires category consent, although another product may use a different approved authority. authorize_delete checks the caller and record before removal. Helper exceptions propagate to the surrounding workflow, which must record partial cleanup and retry it. The fragment has no automatic retry or cross-store transaction.

Code example: Promotion stores the complete active fact. Deletion cleans the listed stores in sequence and reports completion only after those calls succeed.

from dataclasses import replace

def promote_memory(candidate: MemoryFact, user):
    require(candidate.status == "candidate")
    require(candidate.subject_id == user.id)
    require(candidate.source_episode_ids and candidate.provenance)
    require(isinstance(candidate.observed_at, datetime))
    require(user.consent_for(candidate.category))
    approved_use = (candidate.subject_id, candidate.category, candidate.scope)
    verified = validate_candidate(candidate)
    require(verified.status == "candidate")
    require((verified.subject_id, verified.category, verified.scope) == approved_use)
    active = replace(
        verified, status="active", expires_at=retention_policy(verified),
    )
    return memory_store.put(active)

def delete_memory(user, memory_id):
    authorize_delete(user, memory_id)
    memory_store.tombstone(user.id, memory_id)
    derived_indexes.remove(memory_id)
    cache.invalidate(memory_id)
    deletion_audit.record_without_value(user.id, memory_id)
    return {"status": "deleted", "memory_id": memory_id}

Promotion retains the fact’s identity, predicate, value, supporting records, confidence, dates, and intended-use scope. Deletion first marks the primary record deleted, then removes its index entries and cache copy. If index removal fails, the exception stops later calls, so the cache may still contain the value and no completed-deletion audit is written. Reads must check the primary deletion status even during cleanup. The helpers must make repeated cleanup safe, as Section 8.2 explains for idempotent operations. Only after every listed step succeeds does this fragment return deleted. Replicas, backups, exports, and other copies need separate removal or restoration rules. The value-free audit records completion of the listed steps, not proof that every copy everywhere was erased.

Correction rule

When a new fact conflicts with an active fact, check identity, intended use, observation time, and supporting records. If the new value is a correction, mark the old record superseded and make it ineligible for retrieval. Otherwise, retain it only for a distinct valid use. Keep the change visible to the user when appropriate.

A candidate statement may remain in the episodic record while its retention rules permit. Evidence, consent where relevant, confidence, and intended use determine whether it can be promoted to a durable profile fact. Failure to promote it does not authorize keeping it indefinitely.

Conflict resolution compares identity, time, source reliability, and whether the new statement corrects or supplements the old one. For example, a user’s new delivery address supersedes the old address for current deliveries, while the old event may still support a permitted question about a past order. Superseded facts are excluded from current-fact retrieval and stay traceable only while retention permits. Wu and colleagues6 use LongMemEval to test temporal updates, multi-session reasoning, and abstention when the needed information is absent. Evaluation also measures retrieval utility, stale-fact use, correction success, cross-user isolation, and deletion propagation.

11.6 Versioned procedural memory

Correcting a stored fact changes what the application knows about a user or event. Changing a reusable instruction or tool can alter every later action that uses it. The application retrieves semantic records as task context, while versioned prompts, tools, and workflows determine which actions the agent loop can propose. These procedures need tests and release controls before a new version runs.

Prompts, tools, and reusable procedures can change product behavior, so they need a release discipline: review, version, test, deploy, observe, and roll back. This is the book’s adaptation of secure development practice. The NIST Secure Software Development Framework7 supports integrating security into the software lifecycle, but it does not classify prompts or tools as code.

A reusable workflow encoded as a tool or instruction asset saves the agent from re-deriving behavior. It can also preserve a mistake indefinitely. The application therefore records the procedure version with every run and uses conformance tests. A user preference belongs in semantic memory, while a system rule belongs in procedural memory.

Suppose procedure version 1 creates a refund draft and version 2 adds a new eligibility check. Review compares the rule with the approved policy, tests permitted and denied cases, and releases version 2 only if those checks pass. Each execution records version 2 and its result. If monitoring reveals a regression, the application stops affected work and restores a tested version that still satisfies current policy. Rolling back to an obsolete policy would not be a valid recovery. This is an illustrative release trace, not a measured improvement.

Table 11.4: Memory quality requires capture-to-forgetting controls.
Lifecycle stage Memory question Evidence
Capture What event or candidate fact is worth storing? Eligibility rule, required authority, supporting source, sensitivity label
Promote Should this candidate become a durable fact or procedure? Validation, confidence, review where required, and intended-use scope
Retrieve Which memory is relevant to this task and user? Filters, rank, freshness, authorization, token cost
Use How did memory affect the decision or answer? Injected memory IDs and trace attribution
Correct How are conflicts, user corrections, and stale data handled? Supersession record and regression test
Forget When must data expire or be deleted? Retention schedule, cleanup status, downstream removal, and restore rules

Chapter conclusion

Multi-agent systems add specialized control loops when task requirements or a measured one-agent limitation justify them. Durable memory requires a governed persistence lifecycle with explicit capture, promotion, and deletion, rather than an unlimited conversation dump.


  1. Anthropic. (2025, June 13). How we built our multi-agent research system. https://www.anthropic.com/engineering/multi-agent-research-system. The report describes an internal evaluation and token-use measurements for one vendor’s system. This vendor report does not isolate the effect of adding agents from model composition and token use.↩︎

  2. Python Software Foundation. (n.d.). typing: Support for type hints (Python 3.14 documentation). Retrieved September 28, 2026, from https://docs.python.org/3.14/library/typing.html. Python does not enforce function and variable type annotations at runtime. Literal describes allowed values to type checkers. The handoff function supplies its own runtime checks.↩︎

  3. Sumers, T. R., Yao, S., Narasimhan, K., & Griffiths, T. L. (2024). Cognitive architectures for language agents. Transactions on Machine Learning Research. https://arxiv.org/abs/2309.02427. CoALA organizes memory into working, episodic, semantic, and procedural forms. It does not require four physical stores or one database design.↩︎

  4. National Institute of Standards and Technology. (2024). Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile (NIST AI 600-1). https://doi.org/10.6028/NIST.AI.600-1. The profile supports governance of privacy, information integrity, security, measurement, and lifecycle risk. It does not prescribe one retention schedule or the chapter’s evaluation measures.↩︎

  5. Python Software Foundation. (n.d.). dataclasses: Data classes (Python 3.14 documentation). Retrieved September 28, 2026, from https://docs.python.org/3.14/library/dataclasses.html. Dataclasses generate methods from declared fields but generally do not inspect field types. The chapter’s value checks are application code, not behavior supplied by the decorator.↩︎

  6. Wu, D., Wang, H., Yu, W., Zhang, Y., Chang, K.-W., & Yu, D. (2025). LongMemEval: Benchmarking chat assistants on long-term interactive memory. In International Conference on Learning Representations. https://arxiv.org/abs/2410.10813. The benchmark tests information extraction, multi-session reasoning, temporal reasoning, knowledge updates, and abstention. It does not validate the chapter’s storage schema, conflict rule, or deletion process.↩︎

  7. National Institute of Standards and Technology. (2022). Secure Software Development Framework (SSDF) version 1.1: Recommendations for mitigating the risk of software vulnerabilities (NIST Special Publication 800-218). https://doi.org/10.6028/NIST.SP.800-218. The framework supports security practices across the software lifecycle. Applying that release discipline to prompts, tools, and procedures is the book’s adaptation.↩︎