16 AI incidents: containment and recovery
Incident response distinguishes model error, policy failure, hostile content, account compromise, and infrastructure compromise while preserving evidence and testing recovery.
Suppose responders investigate request req-882, whose worker attempted the unapproved external connection recorded in Chapter 15. The monitoring trace establishes a denied connection and a stopped child process, but does not establish their cause. A hostile document in retrieved context is one possible explanation. Responders have to decide whether the event meets the organization’s incident criteria, preserve the records that show what happened, and stop any continuing harmful activity.
Investigation can start with an observed consequence or a warning of attempted harm, then work backward through identities, records, actions, and changed state. A service that is running again has not necessarily recovered. Recovery requires evidence that the identified failure path is blocked in the restored system.
16.1 Incident triage
An alert about an AI system can point to a poor answer, a broken policy, or a real compromise, and each needs a different response. A security event is an observable occurrence relevant to security. An incident is an event that meets the organization’s criteria for response. Triage assesses whether the event involves ordinary model error, policy violation, hostile-content influence, account or infrastructure compromise, disclosure, or resource abuse. These categories can overlap: a wrong answer that discloses a protected record is also a security failure. Classification matters because each path needs different evidence and containment.
NIST1 distinguishes cybersecurity events from incidents and places validation, categorization, prioritization, and incident declaration in the response process. NIST2 also presents response as part of continuing cybersecurity risk management rather than a separate linear phase. The profile does not supply AI-specific categories or one universal declaration threshold.
In the example, the responder starts from a routed alert and checks the affected asset, observed outcome, identities, time window, system version, and supporting trace. A low-quality answer with no boundary violation remains a model-quality issue. A forbidden disclosure or unauthorized action enters incident handling. The declaration record states confidence and missing evidence. Preservation begins as early as the situation allows. When continued activity could disclose more information or execute more unauthorized actions, responders contain it immediately and record which volatile evidence the action may lose.
Triage is the incident lead’s decision. The lead checks the event against documented criteria and affected assets, and the observed outcomes and available evidence decide whether it is declared, downgraded, or escalated. For all routed events, the record reports declarations, later reclassifications, and time to decision separately. Rapid triage consumes responder capacity and can interrupt service. Missing traces or an incorrect model-error label can hide an incident. When that happens, the case is reopened and newly found evidence is preserved. Any containment change records the reason for reclassification. A correct category does not prove complete scope.
16.2 Incident evidence preservation
Some records that explain an incident disappear when a process stops or a log is overwritten. Evidence preservation keeps records available and trustworthy enough to reconstruct the event. Relevant material can include prompts and outputs with retrieved document versions. Policy decisions, tool traces, model and code identities, process state, network events, credential identifiers and use records, and supplier records may also matter. Secret credential values need collection only when they are necessary to investigate the incident and can be stored with appropriate restrictions. Some data may already be missing or inaccessible, which belongs in the evidence record.
The NIST incident-response profile3 calls for recording response actions and preserving the integrity of incident records and data. NIST forensic guidance4 begins with identifying, labeling, recording, and acquiring relevant data while preserving integrity. The forensic publication predates current AI systems and does not supply an AI evidence list or legal procedure.
When capture can proceed without delaying urgent containment, the responder copies available volatile records first and records the acquisition time with its method. Stable files receive hashes. Access is restricted while original versions are preserved. Provider records are requested through the available contract, while unavailable fields remain explicit. Hashes can reveal later file changes but cannot prove that the original record was complete or truthful. Preserved artifacts provide the basis for timeline and scope analysis.
Example
Concrete preservation record. This incident-response exercise follows the denied external connection from worker job-417 for request req-882. The process monitor stopped the newest child, while the worker remains running. Suppose responders can safely capture the remaining worker state while the network control blocks that external route, before isolation or termination. State already lost with the stopped child cannot be recovered by this later capture.
| Order | Evidence source | Acquisition record | Result |
|---|---|---|---|
| 1 | remaining worker process list, open connections, and retained network events | captured at 10:18 by responder R4 |
stored as volatile-01 |
| 2 | application and policy events for req-882 |
exported for 10:10 to 10:20 | stored as events-01 |
| 3 | retrieved document version | copied from source object DOC-55:v7 |
stored as source-01 |
| 4 | container image description | digest and configuration recorded | stored as runtime-01 |
| 5 | external-model provider trace | requested by provider request ID | unavailable at collection time |
For each stable export, the evidence register identifies the collector, time, method, access restriction, and file hash. It marks the provider trace as missing rather than treating silence as proof that no provider event occurred. The conclusion is that the listed records can support a bounded timeline and integrity check. The shortened identifiers and results are invented for the example, and hashes cannot show that the original logs were complete.
The responder counts a source as preserved only after checking its acquisition method, time, integrity record, access rule, and storage result. Coverage is measured against all sources the incident map requires, with acquired, unavailable, overwritten, and untrusted items kept apart. Preservation costs storage and can retain sensitive content longer than ordinary operations allow. Volatile loss, tampered source logs, or inaccessible supplier records can leave gaps in the record. Later snapshots and requests for external evidence can close some of them. Conclusions based on incomplete or untrusted records remain qualified when the incident account is updated.
16.3 Incident scope and timeline
A first alert may identify only one part of an incident. Oligo’s November 2025 ShadowRay 2.0 report5, introduced in Chapter 9 (Runtime intrusion and resource abuse), described attackers moving payload hosting from GitLab to GitHub after takedown. It also reported more than 230,000 internet-exposed Ray servers, compared with a few thousand in its earlier research. Exposure is not a count of confirmed compromises. The hosting change shows why responders must check for replacement destinations, while the exposure estimate identifies systems that may need investigation rather than proving that every one belongs in this incident’s confirmed scope.
Investigation links the initial entry point to affected identities and accessed records. It then checks executed actions and modified artifacts. Observed harm is recorded separately. A timeline orders evidence from application, identity, retrieval, tool, provider, host, and network sources. Competing explanations remain under consideration until the records distinguish them. Attribution confidence remains separate from confidence about the technical path.
NIST forensic guidance6 separates collection, examination, analysis, and reporting across files, operating systems, networks, and application data. The NIST incident-response profile7 includes root-cause analysis as a response outcome. These sources provide process structure without making attribution certain or defining AI-specific artifacts.
For req-882, the team links the retrieved document version and model request to the worker’s process and network records. It checks for proposed tool calls and policy decisions rather than assuming they exist: Chapter 15 recorded no approved tool action. A denied worker connection can occur outside that tool path. The team also checks whether account compromise or changed worker code could explain the same events. Each conclusion cites supporting and conflicting records. The identities, components, sources, and destinations supported by those records define the initial containment set.
Persistence adds a further check because some AI effects remain after removal of the initial trigger. Examples are poisoned model weights, a contaminated retrieval index, copies in caches and agent memory, altered configuration or code, and exposed credentials. Each requires a different recovery path and leaves different evidence. The investigation therefore tags each affected asset with its persistence type and the remaining spread.
The case lead accepts or rejects each path hypothesis after checking its required events against the preserved timeline. Scope is measured over every system, identity, record, and destination reachable from the supported path, and examined and unresolved items are reported separately. Wider scope takes time and delays restoration. Missing telemetry, false correlation, or a second entry point can leave the scope too small. When new evidence widens the path, containment expands and new sources are collected. The affected conclusions are then checked again against the revised timeline. A coherent path does not establish actor attribution or exclude concurrent activity.
16.4 Incident containment
An active incident may continue to expose data or change records. Recovery actions can also erase evidence of what happened. Containment reduces ongoing harm while preserving enough state for investigation. Available actions include revoking credentials and suspending a tool. The organization can also isolate a worker or block a destination. Freezing an affected data source and narrowing user access address other paths. The sequence depends on which path is active and whether a change would destroy volatile evidence.
The NIST incident-response profile8 calls for incidents to be contained and eradicated. International secure AI guidance9 recommends preparing incident procedures and mechanisms for responsible deployment and operation. These outcome statements do not choose the right containment action or sequence for the current incident.
When evidence supports the hostile-document path, the organization preserves the source and traces where this does not delay necessary containment. It disables the affected connector and any credential that could enable the suspected actions. If a worker compromise remains possible, it isolates that workload and blocks external routes. Each containment action has an owner, expected effect, validation check, and reversal condition. Once the path is stable, recovery can restore known assets and permissions.
Example
Containment sequence for req-882. Suppose the follow-up investigation finds that hostile instructions in DOC-55:v7 caused the worker to attempt an outbound network connection (egress). In the reconstructed path, the document supplied a URL that appeared in the model response, and a background worker used that URL for an automatic fetch without an authorization check. That fetch explains why the network event had no approved tool action. These are additional facts assumed for the exercise, beyond Chapter 15’s monitoring trace. With the existing network denial still in place, the responder records remaining volatile worker state, revokes action credential cred-A as a precaution, disables connector conn-RAG-02, isolates worker job-417, and confirms a gateway block for destination 203.0.113.25. The capture records currently open connections, if any, and retained connection events. It cannot turn the earlier denied attempt into evidence of an open connection. Each action records its owner, time, and validation result, such as a denied tool call or blocked connection attempt.
Conclusion. The capture preserves evidence. Credential revocation, connector suspension, worker isolation, and the gateway block restrict the known paths only when their validation checks confirm the expected denials. An alternate destination or a second compromised credential would remain outside this contained set and needs its own check.
The identity service and connector manager revoke access for the supported path, the network gateway blocks destinations, and the workload controller isolates execution. The measure is the number of confirmed blocked paths divided by the total number of known active paths, with service impact and unresolved routes recorded beside it. Broad containment interrupts legitimate work and may destroy volatile state if applied too early. Unknown credentials, alternate connectors, or a compromised control plane can bypass local action. The response then widens isolation, rotates related access, and replaces affected control components before each path is validated again. Blocking the known routes is not proof that the incident has ended.
16.5 Recovery and return to service
Service availability alone cannot establish recovery if the same disclosure or action remains possible. Recovery restores a known state and shows that the incident path no longer works. Affected state may include application code and model packages. System prompts, indexes, caches, agent memory, credentials, tool schemas, and access policy may also require restoration.
The recovery owner can restore an earlier accepted configuration only when it remains available and under the organization’s control. If restoration is unavailable, the affected AI path stays disabled. A fallback may continue work only if its own current access, data, and release-scope checks pass. If no permitted fallback is available, affected outputs are withheld. Required recovery checks must pass before full operation resumes. A failed or unavailable check keeps the affected path restricted or stopped.
The NIST incident-response profile10 calls for checking restoration assets before use, fixing the incident’s root causes before production use, and verifying restored systems. It also calls for confirmation of normal operation before recovery is declared complete. The NIST Generative AI Profile11 recommends post-incident review and practice of response plans. These sources leave AI-specific checks for indexes, caches, memory, and authorization to the recovery design.
In the example, rebuilding the index from approved source versions and clearing affected caches removes known hostile copies. The repair also removes the background worker’s automatic fetch of URLs from model output. Those URLs remain text for display, with automatic link previews disabled. Any required network operation must use the application’s authorized service path. The replacement worker resumes approved retrieval and inference through conn-RAG-02, using replacement credentials limited to those services. Credential cred-A stays revoked. The ban on fetching output URLs and the gateway’s outbound restrictions remain part of the intended operating configuration. Removing the hostile document alone would leave the original cause available to the next document.
Example
Retest before return to service. In an isolated test environment, the team loads the replacement worker, reviewed configuration, rebuilt index, cleared cache, and replacement service credentials planned for return to service. Approved retrieval and inference are enabled. It replays the preserved DOC-55:v7 case and a variant with a different destination under test-team control. The team verifies that output processing handles both cases. Each URL must remain inert: the worker records no fetch request, and network monitoring records no outbound attempt from that output. A gateway denial would reveal that the worker still attempted the prohibited fetch and would fail this repair test, even though it contained the connection. A check that cannot observe the worker’s attempted requests remains unassessed.
A replay in which the model produces no URL cannot by itself test the removed fetch behavior. The team also submits the preserved response and its changed-destination variant directly to the repaired output handler and checks for attempted requests.
The same configuration runs twelve clean tasks, including retrieval, inference, and display of ordinary answers containing URLs. Assume the hostile cases and direct output tests produce no fetch attempt and ten clean tasks pass, while two fail with a permission error. This supports closure of the tested fetch path, but intended work still fails. The team investigates whether the replacement permissions caused the error, fixes the permitted service access, and reruns the affected security and clean cases before release. The permanent restrictions remain in force during that retest.
Conclusion. Return to service requires the original and changed-destination cases to fail at the repaired fetch path while intended work succeeds in the configuration being restored. If safety still depends only on the temporary connector suspension, worker isolation, or block of the original destination, the outcome is continued containment or restricted operation. It does not establish recovery of the enabled service.
Return to service is the recovery owner’s decision. It follows a check of every affected asset, original failure case, clean task, identity, destination, and remaining evidence gap. Each required check records its expected outcome and whether that outcome passed, failed, or could not be assessed. A prohibited action correctly blocked is a pass. A test that could not run is unavailable, not evidence that the action was blocked. Rebuilds and repeated tests extend downtime and can lose useful cached state. An untracked copy, an unchanged supplier component, or a weak regression case can preserve the failure. If a check fails, the attempted restoration is rolled back and service is narrowed. The missing state is corrected before the gate runs again. A passing gate supports only the state it tested.
16.6 Incident follow-up and closure
Follow-up turns incident evidence into changes to controls, tests, contracts, and training. It also decides which records remain for legal, operational, or research purposes and which data and access must be removed. Service retirement includes revoking identities and disabling connectors. It also ends supplier access and checks known derived stores.
The NIST incident-response profile12 integrates lessons and improvement into risk management after response and recovery. NIST media-sanitization guidance13 separates the sanitization decision from verification and validation of the result. Media sanitization alone cannot prove deletion from hosted-model parameters, supplier backups, indexes, or distributed caches.
The organization assigns each corrective action to a responsible owner and links it to a test or contract result. Retention and notification follow current sector and jurisdiction rules, which need their own legal review. Supplier exit evidence records what the organization can verify and what remains outside its visibility. The incident record also shows why harmful activity does not by itself prove that AI caused or improved it. Chapter 17 (AI-assisted attacks) turns from an AI system under attack to an attacker using AI. It keeps the same separation of observation, inference, and unsupported attribution.
At closure, the closure owner chooses retention, deletion, retirement, or reopening. The choice rests on corrective-action evidence, access revocation, supplier response, and applicable obligations. The closure check covers every agreed follow-up action and every service identity or connector in the retirement record. Longer retention raises storage and privacy cost, while early deletion can weaken later investigation. An unlisted supplier copy or forgotten credential can bypass retirement. The hold is then extended when justified, remaining access is revoked, and missing evidence is requested before closure checks run again. Administrative closure does not prove deletion from systems outside the organization’s visibility.
Note
Chapter checkpoint. Follow-up evidence links a hostile document to the attempted external connection for request req-882. The worker has now been stopped, and the provider trace is unavailable. What response record supports containment and recovery, including any evidence that may already be lost?
Answer. Preserve any worker capture taken before termination and the surviving application, retrieval, policy, tool, runtime, identity, and network records. Identify volatile process or connection state lost when the worker stopped. A later snapshot cannot recover it. Record acquisition details, hashes for stable exports, access limits, and the missing provider evidence. Use the available path evidence to guide credential revocation and worker isolation. Urgent containment need not wait for a complete reconstruction. Rebuild the index from approved document versions, clear affected caches, rotate exposed credentials, and remove automatic fetching of model-output URLs. With approved retrieval and inference restored, replay the original and changed-destination cases and confirm no fetch attempt before the gateway, then test clean work. Recovery is supported only for the tested configuration with its permanent restrictions recorded. Missing provider evidence remains a limit on scope and attribution.
Supplier evidence, logical deletion, and jurisdiction-specific notification duties still require case-specific review.
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone (2025), Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3, section 1, printed p. 2, and Table 3, DE.AE-08 and RS.MA-02-RS.MA-03, printed pp. 26-27, source.↩︎
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone (2025), Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3, section 2.1, printed pp. 4-6, source.↩︎
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone (2025), Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3, Table 3, RS.AN-06 and RS.AN-07, printed p. 29, source.↩︎
Karen Kent, Suzanne Chevalier, Tim Grance, and Hung Dang (2006), Guide to Integrating Forensic Techniques into Incident Response, NIST SP 800-86, sections 3 and 3.1, especially 3.1.2, “Collecting the Data,” source.↩︎
Oligo Security (2025), “ShadowRay 2.0: Active Global Campaign Hijacks Ray AI Infrastructure Into Self-Propagating Botnet,” Oligo Security blog, November 18, source. This is the discovering company’s own analysis. Its exposed-server count does not establish the number of compromised servers, and the guide does not independently validate its start date or actor attribution.↩︎
Karen Kent, Suzanne Chevalier, Tim Grance, and Hung Dang (2006), Guide to Integrating Forensic Techniques into Incident Response, NIST SP 800-86, sections 3.1-3.4 and 4-7, source.↩︎
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone (2025), Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3, Table 3, RS.AN-03, printed pp. 28-29, source.↩︎
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone (2025), Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3, Table 3, RS.MI-01 and RS.MI-02, printed pp. 32-33, source.↩︎
UK National Cyber Security Centre, CISA, and international partners (2023), Guidelines for Secure AI System Development, version 1.0, November 27, section 3, incident management and responsible release, and section 4, secure operation and maintenance, pp. 14-17, source.↩︎
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone (2025), Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3, Table 3, RC.RP-03 through RC.RP-06, especially RC.RP-05 R1-R2 on cause repair and restoration checks before production use, printed p. 34, source.↩︎
National Institute of Standards and Technology (2024), Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, action MG-4.2-002, printed p. 45, source.↩︎
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone (2025), Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3, section 2.1, printed pp. 4-6, and Table 2, ID.IM outcomes, source.↩︎
Ramaswamy Chandramouli and Eric A. Hibbard (2025), Guidelines for Media Sanitization, NIST SP 800-88 Rev. 2, section 4.5, “Sanitization Assurance,” especially sections 4.5.1-4.5.2, printed pp. 24-25, source.↩︎