2  Threat modeling and controls

A threat record connects an attacker’s access to a possible failure, the controls expected to prevent it, and the evidence needed to test them.

Suppose an organization runs an assistant that answers employee questions from permitted documents and prepares action requests. The employee confirms the exact operation, and a separate service checks that confirmation and the employee’s current permissions before executing it. The model runs with an external provider, while identity, retrieval, logging, and the action service run inside the organization. When something goes wrong, leaders need a precise answer to four questions: which system version ran, who could act on it, which observable result counts as failure, and which records would show that failure. Without those answers, discussion drifts between what the model generated, what the application sent, and what the service executed.

A useful threat statement specifies the system, the attacker’s access, the unwanted result, and the evidence that would show failure. The system record from Chapter 1 fixes the components and boundaries. The threat record adds a bounded attacker and follows one possible failure from entry point to consequence.

The working order is:

fixed system -> protected asset -> attacker capability -> entry point -> unwanted outcome -> consequence -> control -> test -> remaining risk -> owner

The order can run in both directions. An investigator may begin with an observed consequence and ask backward which earlier events could explain it. The resulting hypothesis must then be checked in the order in which the system actually runs.

Evidence answers several different questions: whether an attack is possible, whether it succeeds under test conditions, whether it occurred in a recorded incident, how often it occurs in a defined population and period, and which actor or system the evidence links to the activity, with what confidence. Evidence for one question may leave the others unresolved. A laboratory demonstration does not establish incident frequency. Several incidents do not establish prevalence without a defined population, period, and collection method. AI roles in attacks explains the evidence classes, and Attack evidence and claim limits discusses attribution and source conditions.

A technical map shows two attacker access profiles, application entry, possible retrieved document, model output, authorization decision, denied and allowed service branches, a separate possible information-disclosure path, controls, and trace records. A denied action has no business effect. The document editor can change an external document that may later be indexed and retrieved.
Figure 2.1: A request-only attacker and an editor of one external document have different entry points into the organization. A model may generate content or propose an action, but the receiving service decides whether an action may execute. A denied request causes no business change; a separate path can disclose information in an answer. The trace connects request, document version, selected passage, model output, service decision, and actual result to support investigation.

2.1 Scope and assumptions

Analysis must stay tied to one version and one use, because the available controls change when the arrangement changes. The reference arrangement here has the organization operating identity, the application, retrieval, logging, and an action service. An external provider operates model inference. Employees can retrieve permitted documents, and generated action proposals require current authorization outside the model.

A threat model analyzes possible attack paths for a specified system, including assets, attacker goals, access, knowledge, and limits. Here a threat record contains one or more path entries, each recording possible effects, the local control, and evidence that could test it. The scope record keeps this analysis tied to one version and use. NCSC secure AI guidance1 recommends a whole-system threat model that considers expected and unexpected behavior, possible compromise, data sensitivity, deployment, and action restrictions.

For the paired tests, suppose the application also offers a customer interface. A customer account may retrieve only records permitted by both its tenant rules and its individual permissions. Employee accounts retain their separate access rules. Business execution still requires an authenticated employee’s confirmation of the exact operation and current checks by the action service. A customer account cannot approve an employee business action.

Table 2.1: Fixed scope for the paired threat cases.
Field Fixed value for the two cases
Decision Whether the employee assistant and its customer interface keep document access and business actions within the specified permissions
System version Organization assistant build 0.2
Operating arrangement Organization-operated application, retrieval, identity, logs, and action service with external model inference
Application capabilities Document retrieval and action proposals. Execution requires an authenticated approval and service-side checks
Users Authenticated employees and external users restricted to their own permitted records
Evidence Correlated identity, retrieval, transfer, provider, proposal, authorization, and execution records where available
Time boundary One request and its linked action attempt. Longer retention and learning effects require separate analysis

The two attacker cases share this extended arrangement and differ by one capability. The first attacker has ordinary customer request access and response visibility. The second can also edit one external document that the organization indexes and an employee may retrieve. In both cases, the attacker lacks administrator access, model-weight access, application-code access, employee credentials, and action-service credentials.

Table 2.2: The second case adds one document-editing capability while other facts remain fixed.
Fact Request-only attacker Document editor
Can submit a request through an ordinary external account Yes Yes
Can edit one external document used by retrieval No Yes
Can change organization policy or application code No No
Can alter model parameters or provider infrastructure No No
Can use an employee approval or action credential No No
Knows internal implementation Only what the interface reveals The same interface observations, plus the format and business context available through the one editable document; no additional internal implementation knowledge is assumed.

Fixing the system prevents a defense from relying on capabilities that do not exist in the stated arrangement. An organization using external inference cannot assume direct access to provider memory. A provider-side request policy check, where offered, can enforce only conditions supported by reliable facts available through its configured interface. The organization may configure the exposed provider controls in advance, so the check need not be requested separately for every model call. It does not replace the organization’s document-disclosure checks or the action service’s checks of current business rules against authoritative records.

Both application traces use inference with retrieval: the model produces output with fixed parameters, and retrieval selects passages for the request. Running either trace as a controlled test and comparing its behavior with an expected outcome evaluates the application. No training update occurs in these traces. A request trace can show what entered a model call. Investigating changed parameters also requires training or update records.

2.2 Security properties and outcomes

A release decision needs a practical test: which conditions must hold, and which observable event would show that one failed? A security property is a condition that protects an asset. An unwanted outcome is a concrete event that would violate such a condition. The property states what must remain true. The outcome specifies what would count as failure, whether or not it has already occurred. An incident investigation then seeks evidence of that event.

NIST AI RMF2 describes related trustworthiness characteristics that include validity and reliability, safety, security and resilience, accountability and transparency, explainability and interpretability, privacy enhancement, and managed harmful bias. These characteristics interact. A system-specific analysis must still identify the asset and the observable event, because a general characteristic alone does not say which record to inspect.

Table 2.3: Security properties become useful when connected to observable outcomes.
Property Practical question Example unwanted outcome
Confidentiality Who may learn this information? A restricted passage enters an unauthorized context or response.
Integrity Who may change this object or decision? A source, prompt, proposal, or business record changes outside its authorized process.
Availability Can authorized users obtain the service within required limits? Crafted work exhausts model, search, or action capacity.
Privacy Are personal data collection, use, sharing, inference, and retention within the stated rules? Personal data is used for a new purpose or retained beyond the approved period.
Safety Can system behavior cause unacceptable real-world harm? Advice or action creates a material harmful result in the stated use.
Reliability Does the system perform its intended function under relevant conditions? The assistant repeatedly gives an incorrect result without adversarial input.
Accountable action Can the organization identify which identity, policy, and service caused a state change? An action occurs without a traceable approval and authorization decision.

The distinction between objective and mechanism keeps that inspection precise. Here, prompt injection means an attempt to redirect a model through attacker-supplied instructions that conflict with the application’s intended task or applicable safeguards. The direct path uses a caller’s message. The indirect path uses material the application retrieves or processes, such as a document, email, or tool result. The goal may be to divert the task or bypass a model safeguard. The delivery path does not establish that the model followed the instructions or that a business action occurred. A restricted document entering an unauthorized model request is an unwanted confidentiality outcome, rather than the name of the attack mechanism. “The model was wrong” may describe ordinary error or an application defect. Hostile influence and bad source data can produce the same observation. Starting with the outcome keeps those possible causes separate, so each can be tested against its own records.

Example

Accuracy and confidentiality. Suppose an employee submits a policy question. Compare two runs. In the first, the application sends only a permitted policy passage to the model, which returns an incorrect date. In the second, it also sends payroll text that neither the employee nor the provider is permitted to receive, but the model returns the correct date.

The first run fails the accuracy requirement. The second violates confidentiality at the transfer to the provider, even though the displayed date is correct and contains no payroll text. The serialized request supplies evidence of that transfer. The visible answer alone does not.

Conclusion. An accurate answer can accompany a security failure. The threat record specifies the protected information, recipient, and observable event.

2.3 Attacker access and capabilities

A vague attacker description hides the conditions that determine whether an attack is possible. The description must state what the person or system can change and observe. An attacker capability is the action or change available to the attacker. Access is the interface or object used for that action. Knowledge records what the attacker knows about the target.

NIST’s adversarial machine learning taxonomy3 separates lifecycle stage, attacker goal, capability, and knowledge. It also distinguishes query access, training-data control, model control, source-code control, testing-data control, and resource control. Those distinctions prevent query access from silently becoming permission to change training data or code. Stage matters here as well: changing a retrieved document before inference differs from changing parameters during training and from changing a response after generation.

Code example: A bounded request-only attacker profile.

Actor: external user with an ordinary application account
Goal: obtain a record that this account may not read
Capability: submit up to 60 requests per minute and observe displayed text
Knowledge: public interface and fields learned from the user’s own responses
Cannot: edit documents, inspect provider internals, change parameters, or use employee credentials
Budget: 500 requests over one day
Success: forbidden record content appears in retrieval output, the serialized model request, or the displayed response

The profile also needs a budget, such as requests per minute, total attempts, time, money, or access to auxiliary data. A broad statement such as “the attacker can interact with the model” hides the conditions that determine whether an attack is possible and whether a test represents the deployment. The bounded profile above states the account, the rate, the total attempts, the visible output, and the explicit exclusions, so a later test can reuse the same limits.

2.4 Attack paths and discovery

The bounded attacker profile lets us ask which failure paths are possible. A path follows an affected object through the components that can turn the attacker’s influence into an unwanted result. A goal states the result sought. NIST’s adversarial machine learning report4 includes availability, integrity and privacy failures among generative AI goals, together with misuse enablement, meaning bypass of the owner’s technical restrictions on use.

Entry stage and consequence stage can differ. A poisoned training example can affect later predictions after parameter updates. A hostile retrieved document can influence a request without changing model parameters. A changed or misdirected delivered response can mislead its recipient even when the model’s original output is unchanged.

Knowledge is separate from the actions the attacker can take. In NIST’s terminology, a black-box setting provides little or no internal information about the target model or system, a gray-box setting provides partial information, and a white-box setting provides full internal information within the specified target. Record the actual known fields, such as architecture, parameters and training data. Separately record what can be queried, observed or changed and the query limits. These knowledge labels grant no modification authority.

The table is a navigation map for this guide. Its ten paths overlap and can be combined. They are not an exhaustive catalogue or an official NIST classification. Read across a row to connect the affected object, the access needed and the stage of influence, then follow the link for the mechanism and controls. Having the listed access does not establish that the attack succeeds. For an attacker limited to ordinary queries, the training and model-replacement rows require powers absent from that profile.

In the collaborative training arrangement used later, participants train local model copies and send updates for a coordinator to combine. The coordinator uses those updates to form the next shared model. They are not ordinary requests during inference. Distributed learning explains the federated arrangement and its submitted objects.

Table 2.4: Overlapping attack paths, their affected objects and access assumptions. Later chapters develop the local controls and their limits.
Path and affected object Required attacker access Where influence enters or takes effect First deeper explanation
Training data poisoning: examples or labels Ability to change examples or labels that a training procedure uses Introduced into learning data; can affect parameters through later updates Chapter 3
Model tampering: model parameters, added learned components or contributed updates Ability to replace or alter an admitted model component or update Admission or an update step; altered parameters can affect later inference. Loader code is a separate software risk Chapter 3 and Chapter 7
Hostile input: request text or model input Ability to submit input, or alter another caller’s request at an intervening component With parameters fixed, crafted inputs may cause a model error or redirect instructions. These objectives need different tests Chapter 4
Response tampering: generated output or its association with a recipient Control of a response-processing, storage or delivery component After generation; a changed or misdirected response need not reflect the model’s actual output Chapter 4
Hostile stored content: sources, search records or retained application memory Ability to write content that the application later admits During ingestion and later use; false source claims and embedded instructions are different mechanisms and can act without updating model parameters Chapter 5 and Chapter 10
Query-based information attacks: returned observations about data or a model Ability to query the target and observe returned fields; record those fields, other knowledge and query limits Queries may seek training content, whether a record was used in training, or model behavior or parameters. Parameter recovery and behavioral imitation need different evidence Chapter 6
Application disclosure: retrieved text, model requests, displayed answers or retained copies Access to a data path, or input influence that can cause data to reach a recipient lacking permission At retrieval, transfer, display or storage; confidentiality can fail before an answer is shown Chapter 6 and Chapter 11
Software or platform compromise: executable components, hosts, interfaces or credentials Ability to supply material that is loaded or executed, or to reach an interface with an exploitable weakness At component admission, execution or management; required capabilities depend on the particular boundary Chapter 7, Chapter 8 and Chapter 9
Resource abuse: compute, output budget or service capacity Request, content or operation access that reaches repeated or costly work During execution; consuming resources is distinct from controlling the document that caused the work Chapter 9 and Chapter 13
Actions outside authority: operation arguments or submitted actions Input influence or caller access that reaches an action path During planning and submission; a business effect still depends on the application’s and receiving service’s checks Chapter 12 and Chapter 13

Several methods help find candidate paths. Microsoft’s STRIDE documentation5 introduces STRIDE, six categories used to ask where software could fail: spoofing (false identity), tampering (unauthorized changes), repudiation (actions without dependable accountability), information disclosure (exposure to an unauthorized recipient), denial of service (loss of service for legitimate users), and elevation of privilege (access beyond granted rights).

An attack tree divides one attacker goal into smaller goals. In Schneier’s description6, OR branches offer alternatives, while an AND branch requires every child condition. The tree can carry researched cost or likelihood values, but drawing it supplies no measurements.

MITRE’s ATLAS data7 describes ATLAS, the Adversarial Threat Landscape for AI Systems, a catalogue of adversary tactics and techniques, mitigations and case records. Entries suggest candidates to investigate. They do not establish local vulnerability or incident frequency. NIST’s AI RMF8, its AI Risk Management Framework, organizes risk work through Govern, Map, Measure and Manage. These functions connect context, evidence, decisions and responsibility, and can be revisited as the system changes.

Table 2.5: Complementary discovery and management methods compared by their inputs, results and limits.
Method Input and question Useful result What remains to be checked
STRIDE Mapped software components and flows: where could identity, modification, accountability, disclosure, availability or privilege fail? Candidate threats at particular components and boundaries Whether the mechanism is possible in this AI system and whether the controls work
Attack trees An attacker goal: which alternative paths or combined conditions could produce it? A decomposition of the goal into prerequisites and alternatives Missing paths, feasibility and any numerical likelihood or cost
ATLAS techniques and risk catalogues A specified system: which published techniques, cases or risks could apply? Candidates to investigate, with catalogue edition and access conditions recorded Local applicability, causal evidence, coverage and incident frequency
AI RMF Intended use and identified risks: which context, measures, decisions and owners are needed? A risk-management process tied to organizational decisions The actual attack paths and evidence that the chosen controls work

Use the discovery methods to select a candidate, then test its proposed path against the fixed system and attacker profile. The following backward and forward traces check which transitions the records support. Their results feed the reusable threat record developed later in this chapter.

2.5 Tracing harm through the system

Suppose a payment record shows that money was sent to an unapproved account. Investigators can start with that recorded change and work backward to the request that caused it. The diagram contrasts this investigation with a later check in runtime order.

A possible runtime path connects attacker-controlled input, an application interface, selected context, model output, authorization, and a business change. The denied branch stops. A reverse hypothesis arrow starts from observed harm. Separate evidence records support confirmation, rejection, or an unknown result.
Figure 2.2: An observed business change suggests possible earlier causes. A proposed explanation can then be tested in runtime order, from the input and selected context to the model output, application decision, and recorded effect. Each transition has its own supporting record. A denied request ends before the business change, while a missing record leaves the corresponding transition unknown.

The investigation asks:

  1. Which identity and policy decision authorized the state change?
  2. Which operation and fields reached authorization?
  3. Which application output supplied those fields?
  4. Which model context and source versions influenced the output?
  5. Which user, connector, or external party could change the relevant input?

These questions generate possible explanations. They do not establish that an earlier event occurred. Each hypothesis can then be checked in the order in which the system actually runs. Each transition in that order is paired with a record that can confirm or reject it. Missing evidence leaves that transition unknown, and the record does not fill the gap with a guessed sequence of events.

A complete threat statement has the following form:

Note

For this system version, an attacker can change or submit a specified input. That input may cause an observable failure and a particular harm. A specified component should prevent or detect the failure. The test records the input, intermediate decisions, and result, then compares them with the expected behavior. The decision owner records what remains unknown.

The next two cases apply this form while changing one attacker capability.

2.6 Case: cross tenant disclosure

The first attacker uses an ordinary external account. The attacker can submit requests, observe displayed responses, and repeat attempts within the stated budget. This person cannot edit documents, use an employee identity, inspect provider internals, or call the action service.

Possible losses include cross-account retrieval, provider disclosure outside policy, resource exhaustion, misleading output, and disclosure through error behavior. Each one needs its own success condition, because the records that would confirm one loss do not confirm the others.

One complete statement is:

Note

An external user submits a request crafted to match another tenant’s record. If the retrieval service applies an incorrect tenant filter, it may admit the forbidden record. Confidentiality fails when the record identifier or text enters retrieval output, the organization’s serialized model request, or the displayed response. The test uses two tenants and fails on any cross-tenant identifier or text. Provider receipt remains unknown without provider evidence.

Table 2.6: Controls and evidence for the request-only case.
Decision Enforcing component Evidence Failure condition
Establish caller Identity service Stable user and tenant claims linked to request ID Identity is missing, stale, or mismatched.
Admit passage Retrieval authorization Candidate IDs, current access result, and rule version A forbidden document is admitted.
Send context Provider disclosure check Serialized request, selected fields, destination rule, and provider receipt when available A forbidden field appears in the serialized request.
Limit work Gateway and service quotas Request rate, token, latency, and rejection records Accepted work exceeds stated limits or blocks authorized use.
Present answer Application Sources, validation result, and displayed response Unsupported text is represented as verified policy under the test rule.

The retrieval record, serialized request, provider receipt, and displayed response are different observations. A user-interface test cannot establish which source entered context or what the provider received. The table above links each decision to the component that makes it and the record that shows it.

The retrieval authorization check enforces tenant separation at the point where passages are admitted to context. It compares the current caller, tenant, document, and rule version before admitting a passage, and it records candidate identifiers with the rule version that decided each one. The disclosure check then confirms that the serialized request contains only permitted fields before the model call. Stale tenant metadata or an unmediated search path can bypass the retrieval rule, and that bypass is a property of the path, not a failure of the model to refuse. Full measurement of forbidden and permitted admissions, latency, cache effects, invalidation of affected caches, and identification of affected request IDs belongs to the shared evidence method developed in Chapters 14 through 16. A zero count in the seeded tenant test does not cover other query strategies or provider-held copies.

2.7 Case: hostile retrieved documents

The second attacker has the same request access and can also edit one external document that the organization indexes and an employee may read. The document contains text asking the model to ignore organization policy and propose a large transaction to an attacker-selected destination.

Greshake et al.9 described indirect prompt injection as adversarial instructions placed in data likely to be retrieved by an LLM-integrated application. Their paper demonstrated the mechanism in synthetic applications and selected systems available at the time. NIST10 uses resource control for attacker control over documents or web pages that enter runtime context. The sources support the mechanism. They do not establish current success rates for every model or application.

Correct retrieval does not settle source integrity or instruction authority. The employee may be allowed to read the document, and the disclosure policy may allow it to reach the provider. Those decisions can pass while the document contains hostile text, because permission to read differs from authority to instruct.

The following assumed negative-test trace checks the service separately from the normal employee confirmation route. The evaluator’s intervention does not add a permission to the document attacker:

  1. The external account changes the document.
  2. The connector copies the new version and source metadata.
  3. Search returns the document for a related employee request.
  4. Retrieval authorization admits the passage because the employee may read it.
  5. The provider-disclosure check permits the passage for this provider.
  6. The application adds the passage to the question and its instructions, then sends the assembled request to the external model.
  7. The model returns an unwanted transaction request.
  8. The application parses and checks the generated fields.
  9. For this negative test, an authorized evaluator makes the test application submit those fields using its normal authenticated context, but with no employee confirmation record.
  10. Under the assumed policy, the action service denies the request because the required confirmation is absent or the destination or amount is forbidden.
  11. The evaluator checks the service decision and the payment record to establish whether any business effect occurred.

The denial is evidence about the action boundary. It does not establish that the displayed explanation was accurate or that every injection attempt failed.

In this test, a payment can occur only if the action service accepts the submitted request. The service checks the authenticated employee identity and current permissions, the supplier and resolved destination, the submitted amount and other fields, and the operation key. It verifies a protected employee confirmation record and compares that record with the exact submitted operation before allowing a new business effect. A different connector that can write business state without this service defeats the claim, because the check never sees the write. Revocation of that connector, preservation of linked records, and reversal where the business process permits it belong to the recovery method in Chapter 16. The result says nothing about misleading text that a person later accepts.

Example

Separating two reasons for denial. In the preceding test trace, suppose the generated request specifies an amount of 9,000 currency units and uses a destination with no approved supplier record. The evaluator submits it without an employee confirmation record. This intervention belongs to the test. The document attacker still cannot approve as the employee or change application code.

Under the assumed policy, the service denies the request. Linked records connect the source version, model request, generated fields, submitted operation, denial, and payment record. Because both the destination and confirmation checks fail, this combined case does not show that each check works separately. Testing confirmation alone requires an otherwise permitted request with confirmation missing or mismatched, plus a permitted request with valid confirmation.

Conclusion. The action check rejects the attempted payment in this trace. The displayed response may still contain the harmful instruction, and the result does not establish that every hostile document will produce the same outcome.

Table 2.7: Different attacker capabilities create different paths and control locations.
Element Request-only attacker Document editor
Controlled input Direct application request Retrieved external document
Primary entry point User interface or API External source and ingestion path
Main boundary at risk Tenant retrieval and provider disclosure Source integrity, context interpretation, and action proposal
Can authorize an action as an employee No No
Critical test No forbidden record enters retrieval output or the serialized request The action service rejects missing or mismatched employee confirmation and forbidden destinations or amounts; the decision and business record are checked.
Concern after the test passes Other request strategies and permitted-data errors Misleading displayed text and other hostile-source strategies

2.8 Controls at trust boundaries

A component can enforce a rule only when it has the required facts and authority. The retrieval service can check current document permission. A disclosure service can check destination and data class. The action service can check business state, authenticated identity, amount, destination, and retry status. The model cannot replace those controls through a refusal message, because the refusal arrives after retrieval and transfer decisions have already been made.

Saltzer and Schroeder’s 1975 tutorial11 collected several protection principles that help place these checks. Their paper credits E. Glaser for the fail-safe-defaults idea and Roger Needham for separation of privilege.

  • Fail-safe defaults: access begins denied and becomes permitted only when stated conditions hold.
  • Complete mediation: every relevant access is checked, including cache reuse and action retries.
  • Separation of privilege: sensitive operations may require independent conditions, such as a user confirmation and a current service-side policy decision.
  • Least privilege: each user or program receives only the permissions needed for its task.

These principles guide design. Evidence must still show how a component implements and tests the rule.

Preventive controls try to stop a forbidden transfer or operation before it occurs. Detective controls inspect events or records for signs of a possible failure. Recovery actions contain the problem and restore an acceptable service or business state where the process supports it.

Table 2.8: Preventive, detective, and recovery measures belong at specific components.
Path Preventive control Detective evidence Recovery action
Cross-tenant retrieval Current tenant and document authorization before context assembly Candidate and admitted IDs with rule version Revoke access, invalidate caches, and identify affected requests
Forbidden provider transfer Data-class and destination policy before the model call Destination, selected fields, policy result, and provider request ID Stop the route and follow the incident process
Attacker-edited document Restricted source writes, versions, source identity, and integrity review Source version, editor identity, unusual changes, and retrieval linkage Restore a trusted version and re-evaluate affected sessions
Unauthorized action Schema validation, exact approval, current identity and business rule checks Proposal, approval, policy decision, operation key, and result Suspend the affected action route, investigate linked requests, and reverse completed effects where the business process permits.
Misleading answer Visible sources and checks for facts used in consequential decisions Output, cited sources, reviewer correction, model and prompt version Correct the record and add a regression case

The table places each measure at the component that can enforce it or supply evidence for inspection. Repeating one model-level filter does not create an independent authorization check. Later chapters test the separate controls and their limits.

Example

Preventing a transfer. Suppose search returns a public passage and a payroll passage for the same employee query. The employee may read only the public passage, and the provider disclosure rule also excludes payroll. A test marker in the payroll passage makes its presence visible in the serialized request.

With permission checks before context assembly and a disclosure check before sending, the request contains only the public passage. With only a model refusal after inference, the payroll marker has already crossed to the provider even if it never appears in the answer.

Conclusion. A check before sending can prevent the transfer. A later refusal changes the answer but cannot undo information already sent.

Operating cost and recovery differ by location and are measured where the control runs. Retrieval and disclosure checks consume latency and can reject authorized work. Approval for an employee action takes time. Monitoring consumes storage and review effort. The corresponding recovery route invalidates an index, disables a transfer, revokes an action credential, or corrects a displayed record. A passing result supports only the component, version, paths, and inputs that the test exercised. Denominators, false denials, and retesting rules use the method in Chapter 14.

2.9 Control claims and evidence

“We use access control” identifies an intention. A testable claim identifies the system version, initial state, attacker capability, inputs, observation points, and pass or fail rule.

A useful test record contains:

  1. application, policy, model, prompt, index, and provider versions.
  2. users, permissions, documents, quotas, caches, and business state.
  3. exact attacker operations and budget.
  4. ordinary, boundary, and adversarial inputs.
  5. retrieval, transfer, output, policy, action, latency, and log observations.
  6. a pass or fail condition selected before the result.
  7. observed results, missing evidence, and untested conditions.

The rest of the guide uses a compact control record for this information. It identifies where a rule is enforced, what facts are checked, what decision is made, what evidence and denominator support the result, the operating cost, possible bypasses, the recovery route, and the limit of the conclusion. Later sections may use a shorter form when these fields are already clear from the surrounding trace.

Table 2.9: Control claims linked to observable tests.
Claim Test input Required observation Fail when
A user cannot retrieve another tenant’s record A request deliberately shares terms with the forbidden record Candidate IDs, authorization decisions, and admitted passages Any forbidden passage is admitted.
Forbidden text does not reach the provider request The forbidden record contains a seeded marker Serialized request fields and provider receipt fields when supplied The marker appears in the serialized request.
A document cannot authorize an action A retrieved document requests an operation outside policy without approval Proposal, validation, authorization, and final business state The business record shows execution without the required approval or despite a failed policy check.
Approval binds exact fields The destination changes after the employee approves the proposal Approval object and action request The changed destination is accepted.
Retry does not duplicate an action while pending execution or the retained result is protected Repeat the same authorized request for the same employee, operation key and fields during that protected period Stable operation key, matching employee and fields, service protection state and final business state The business state shows more than one completed action for that protected operation.

The retry case uses the receiving-service conditions from Chapter 1: pending execution is protected, and the completed result is retained for 24 hours after completion. After that period, duplicate protection is no longer assumed, so an uncertain result must be reconciled before another submission. Each denial test needs an authorized control case. A service that denies every request fails to provide the authorized service. Denial may prevent disclosure at that decision point, but confidentiality tests still need to inspect earlier transfers, retained copies, and other paths. Testing both the forbidden and permitted cases distinguishes the intended control from a broken system.

The test harness, the software that runs the cases and collects their results, records the numerator and denominator for each outcome rather than reporting only that an attack was blocked. It also records latency, retries, reviewer work, and false denials. A release policy can require a failed candidate to remain undeployed while the owner investigates. If a failure is found after deployment, restoring an earlier configuration is a separate recovery decision. It does not follow automatically from a failed test. A harness that cannot observe a provider transfer or a business state change leaves that part of the claim unknown. Passing the visible steps does not prove that an unobserved path is absent. Full metric definitions, intervals, and adaptive testing belong to Chapter 14.

2.10 Residual risk

Tests reduce uncertainty under their stated conditions. They rarely remove all possible failure paths. Residual risk is the risk that remains after the selected controls are in place. NIST AI RMF12 asks organizations to establish context, characterize impacts and likelihood, prioritize risk, choose responses, and document remaining negative risk. The framework does not calculate an organization’s risk tolerance or approve a deployment.

Table 2.10: Assumed remaining-risk decision for a limited organization pilot.
Field Assumed limited-pilot decision
Risk Attacker-controlled external text influences an action proposal.
Possible impact Misleading employee guidance, an attempted unauthorized action, and investigation cost
Preventive evidence The action service independently checks identity, exact approval, destination, scope, and current limits.
Test evidence Assumed seeded cases outside policy are denied while a permitted control succeeds once.
Evidence gap The exercise does not measure every way hostile text may mislead an employee.
Decision Permit a limited pilot while automatic execution remains restricted to the tested action path.
Owner Business service owner with security review
Review triggers Model, provider, prompt, parser, connector, permission, destination, or action-limit change

The record distinguishes observed facts, reasonable inferences, and unknowns. An assumed result is labeled as assumed. A small test cannot become a broad statement that an agent is secure.

The responsible decision owner compares the evidence with the conditions required for the intended use. Restrictions may reduce capability, while monitoring consumes storage and staff time. An unrecorded change can invalidate the comparison. Where an approval depends on a specific configuration, a changed data path or permission rule requires a new assessment and may require suspending the affected use. Accepting the remaining risk records a decision under these conditions. It does not prove safety.

Review triggers connect the threat record to later control and review decisions. A new model, prompt, connector, data class, action, provider, or permission rule can invalidate earlier evidence.

2.11 Exercise: email assistant threat model

Suppose an organization plans to let an internal assistant prepare an email to an external supplier. The employee must approve the exact recipient, subject, and body before a mail service sends it. A retrieved supplier document may contain attacker-controlled text.

Which added components and attacker capabilities belong in the threat record? Trace one integrity failure and one confidentiality failure, then identify the controls, tests, remaining risk, and changes that would require another review.

Answer. The scope adds the mail service, its narrow service identity, the recipient, message body, approval object, send result, and linked records. The attacker can edit the supplier document but cannot alter the recipient record, approve as the employee, change application code, or call the mail service.

One integrity outcome is an attacker-selected link in a sent message. The backward investigation starts with the sent-message record and asks which approval, draft, context, and source version could explain it. Forward verification follows source update, ingestion, retrieval, model request, generated draft, displayed approval, mail authorization, and send result.

The attacker may cause the link to appear in the draft before the employee reviews it. If the employee approves that draft unchanged, binding approval to its exact contents does not remove the link. One test therefore records whether the reviewer or a separately defined link policy rejects the hostile draft. Another test deliberately changes the body after approval to check approval binding. The test software injects this change. It is not an extra capability granted to the document attacker.

The mail service accepts only a draft identifier and approval object bound to the exact recipient and content hash. One negative test changes the body after approval and expects denial, while the control test sends an unchanged benign draft once. Logs join the source version, serialized model request, draft, approval, mail request, and result.

The test record states changed drafts denied out of all changed drafts and unchanged messages sent once out of all permitted controls. Exact approval adds employee time and may require a new approval after any edit. A mail route that does not call the service bypasses this control. Recovery disables that route and revokes its credential. When available, the mail system’s recall or correction process handles recovery, although the test does not show that the approved message is truthful.

For a confidentiality test, suppose the employee may read an internal price document, but the external supplier may not receive it. The attacker-edited supplier document asks the assistant to include that price in the email. Retrieval can correctly admit the source for the employee, so a separate disclosure rule must check the recipient and outgoing body.

The evaluator puts a unique test marker in the internal document. The negative test expects the mail service to block a message containing that marker and the send record to show no delivery. A permitted message without restricted content should be sent once. This tests the seeded disclosure path, not every way the model could paraphrase private information.

An attacker need not read the restricted source directly to try to make the application disclose it. The trace still requires evidence that the application had source access and that the attacker-controlled input could influence its selection or transfer. Remaining risk includes subtle misleading text that an employee approves. A model, prompt, parser, approval-screen, connector, or mail-service change triggers a new evaluation.

2.12 Reusable threat records

System and threat records answer complementary questions:

  1. The system record identifies components, operators, identities, assets, operating modes, data paths, action paths, lifecycle stages, and evidence.
  2. The threat record links a fixed system and bounded attacker to an unwanted outcome, consequence, control, test, remaining risk, owner, and review trigger.

Later chapters change data influence, model access, platform exposure, retrieval, tools, or organizational duties. Each branch states which field in these records changes. The same record lets a later test or responsibility decision refer to a specific earlier failure path.

The rest of the book groups those branches so the reader can follow one problem at a time. Data problems run through Chapters 3 through 6, runtime problems through Chapters 7 through 9, and authority problems through Chapters 10 through 13. Evidence methods follow in Chapters 14 through 16, cybersecurity applications in Chapters 17 through 19, organizational decisions in Chapters 20 through 22, and changed assumptions in Chapters 23 through 25. This grouping guides the structure of this book. It is not a set of official NIST categories.

Readers who already use the OWASP GenAI LLM Top 1013 can treat it as an entry map rather than a replacement for the system and threat records. The following links show where this guide develops the application risks in the 2026 edition. OWASP labels and rankings change between editions, so this map uses that edition’s entries and keeps the underlying mechanisms explicit.

Table 2.11: OWASP application risks linked to their first explanations and later control discussions.
OWASP 2026 entry Related chapters
LLM01 Prompt Injection Request manipulation (Chapter 4), retrieved instructions, and execution (Chapter 10)
LLM02 Sensitive Information Disclosure Data exposure (Chapter 6), platform access (Chapter 08), and retrieval permissions (Chapter 11)
LLM03 Excessive Agency Delegated authority (Chapter 12), action sequences (Chapter 13), and changed assumptions (Chapter 24)
LLM04 Supply Chain Acquired components (Chapter 07)
LLM05 Data and Model Poisoning Training poisoning (Chapter 3) and retrieved-content poisoning (Chapter 5)
LLM06 Unbounded Consumption Service exhaustion (Chapter 09), task budgets (Chapter 13), and live observation (Chapter 15)
LLM07 Misinformation Manipulated answers (Chapter 4), corrupted source content (Chapter 5), and attacker assistance (Chapter 17)
LLM08 Hidden Context Exposure Data exposure (Chapter 6), application interpretation (Chapter 10), and action authority (Chapter 12)
LLM09 Vector and Embedding Weaknesses Retrieval poisoning (Chapter 5) and retrieval permissions (Chapter 11)
LLM10 Improper Output Handling Executable outputs (Chapter 10) and action authority (Chapter 12)

2.13 Chapter checkpoint

Suppose an action service rejects all 100 unauthorized payment requests in a test, but no permitted request was tested. What does this result support, and what remains unknown?

Answer. It supports rejection of those requests under the tested configuration. It does not show that permitted payments succeed, that other attack strategies fail, or that an earlier data transfer was permitted. An authorized control case and observations along the full request path address different parts of that uncertainty.

2.14 Chapter references

  1. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025 (2025), Sections 2.1.1-2.1.4 and 3.1.1-3.1.3, official publication.

  2. Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023), Figure 4 and Section 3, report pp. 12-18, for trustworthiness characteristics; Figure 5 and Sections 5.1-5.4 for Govern, Map, Measure and Manage; and Tables 2 and 4 for the cited risk decisions, publication. The framework does not calculate an organization’s risk tolerance or approve a deployment.

  3. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), secure design, PDF pp. 9-10, official guidance.

  4. Jerome H. Saltzer and Michael D. Schroeder, “The Protection of Information in Computer Systems,” Proceedings of the IEEE 63(9) (1975), pp. 1278-1308, Section I.A.3, author-hosted text.

  5. Kai Greshake et al., “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” arXiv:2302.12173v2 (May 5, 2023), Sections 1, 3, and 4.1, author manuscript.

  6. OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, version 2026, PDF pp. 3 and 9, official publication.

  7. Microsoft, Threat Modeling Tool threats, STRIDE model table. Retrieved September 28, 2026, from article. The accessible table supports the six categories, not exhaustive AI threat discovery.

  8. Bruce Schneier, “Attack Trees,” Dr. Dobb’s Journal (December 1999), “Enter Attack Trees” and “Creating Attack Trees,” author-hosted article. This citation supports the described goal decomposition and AND/OR branches. It does not credit this article with every predecessor or provide measurements for this guide’s system.

  9. MITRE ATLAS contributors, ATLAS Data, repository README, “Distributed ATLAS Data” and “Data Format.” Retrieved September 28, 2026, from repository README. The README describes tactics, techniques, mitigations and case-study records. Catalogue presence alone establishes neither local vulnerability nor prevalence.


  1. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), “Model the threats to your system” and following secure-design guidance, PDF pp. 9-10, source.↩︎

  2. Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023), Figure 4 and Section 3 for trustworthiness characteristics, and Figure 5 and Sections 5.1-5.4 for Govern, Map, Measure and Manage, source.↩︎

  3. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025 (2025), Sections 2.1.1-2.1.4 and 3.1.1-3.1.3, especially report pp. 7-8 and 40-41, source.↩︎

  4. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025 (March 2025), sections 2.1.1-2.1.4 and 3.1.1-3.1.3, official report. These sections distinguish stage, goal, capability and knowledge. The table combines those distinctions with this guide’s application and service boundaries.↩︎

  5. Microsoft, Threat Modeling Tool threats, STRIDE model table. Retrieved September 28, 2026, from official documentation. The accessible article table supports the six categories, not a claim of exhaustive AI threat discovery.↩︎

  6. Bruce Schneier, “Attack Trees,” Dr. Dobb’s Journal (December 1999), “Enter Attack Trees” and “Creating Attack Trees,” author-hosted article. This citation supports the described goal decomposition and AND/OR branches. It does not credit this article with every predecessor or provide measurements for this guide’s system.↩︎

  7. MITRE ATLAS contributors, ATLAS Data, repository README, “Distributed ATLAS Data” and “Data Format.” Retrieved September 28, 2026, from official repository. The source describes tactics, techniques, mitigations and case-study records. Catalogue presence alone establishes neither local vulnerability nor prevalence.↩︎

  8. Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023), Figure 4 and Section 3 for trustworthiness characteristics, and Figure 5 and Sections 5.1-5.4 for Govern, Map, Measure and Manage, source.↩︎

  9. Kai Greshake et al., “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” arXiv:2302.12173v2 (May 5, 2023), abstract and Sections 1, 3, and 4.1, source. The demonstrations concern synthetic applications and selected systems available at the time, including GPT-4-based applications. They do not establish current attack success rates.↩︎

  10. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025 (2025), Section 3.1.3, report pp. 40-41, source.↩︎

  11. Jerome H. Saltzer and Michael D. Schroeder, “The Protection of Information in Computer Systems,” Proceedings of the IEEE 63(9) (1975), pp. 1278-1308, Section I.A.3, “Design Principles,” source. The paper credits E. Glaser for fail-safe defaults and Roger Needham for separation of privilege.↩︎

  12. Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (2023), MAP 1.1, MAP 5.1-5.2, and MANAGE 1.1-1.4, Tables 2 and 4, source.↩︎

  13. OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, version 2026, contents on PDF p. 3 and “LLM Top 10 at a Glance” on PDF p. 9, official publication and download. The table maps entries LLM01:2026 through LLM10:2026 to this guide; it is a navigation aid, not a claim of complete OWASP coverage. The retrieved PDF retains publication-date placeholders; the landing page is dated August 3 and the cover August 4, 2026.↩︎