13  Limits on agent execution

Browser sessions and coding workspaces give agents access to files, credentials, and services, so execution limits and checks on each result must match the authorized task.

In July 2025, Jason Lemkin reported that a Replit coding agent deleted a production database during a code freeze, a period when no changes were supposed to be made. The agent then told him a rollback would not work, which was wrong. Fortune1 covered his account on July 23, 2025. The cited coverage reports his account rather than a published forensic investigation. Replit added separation between development and production afterwards.

In that account, the environment permitted deletion of production data despite the code freeze. Checking each tool call, as in Chapter 12 (Agent action authorization), does not settle which files or services an executed action can affect. That depends on the environment. Browser sessions, code workspaces, images and audio, and long action sequences bring state and authority that one tool call may not reveal. Suppose an organization runs a support tool in a browser session or a coding workspace. Its record states the session identity, files, secrets, network access, and permitted operations, and it keeps proposed actions separate from executed ones.

A controller reads a trusted task and separately labeled workspace and media data. Coding review and browser authority form distinct branches. A confined runner limits files, credentials, network, and resources, while state and result checks connect each action to continuation, stopping, or recovery. The media panel shows screenshot, audio, and video inputs passing through extraction and source-labeled context to the model and a proposal.
Figure 13.1: The trusted task remains distinct from workspace text and interpreted media as the controller works across code and browser environments. Generated changes require review, browser actions depend on session and target state, and execution limits restrict access and resource use. The controller checks the resulting state before work continues or recovery begins. The combined action sequence is the behavior evaluated for release. The media panel illustrates an explicit extraction path, discussed in Multimodal prompt injection. Systems that submit media directly to the model may have no separate optical character recognition or transcript record.

13.1 Hostile workspace content

A coding assistant must read repository files and test results to complete its task, but those files can contain instructions the employee never authorized. An untrusted workspace contains files and messages needed for a task but not authorized to direct the agent. Source files, issue text, logs, test output, web pages, and downloads can all contain text that resembles a user instruction. The agent still needs to read that material without letting it replace the employee’s task or expand permitted actions.

Greshake et al.2 showed that external data retrieved during a task can deliver indirect prompt injection to an integrated language-model application. OWASP’s 2025 prompt-injection entry3 includes external websites and files among possible injection inputs. These sources cover selected delivery paths and a community risk model. They do not measure every coding-agent workspace or build system.

The example organization’s coding assistant keeps the employee request in a trusted task record and labels repository, issue, and command output as data. Each proposed action still passes the tool contract and policy checks from Chapter 12. A test places conflicting text in an issue and verifies that the agent reports it without executing its request. This input boundary sets the review rules for generated software and dependencies.

The workspace reader records source labels; the action gateway checks the trusted task, current identity, proposed operation, and destination. A blocked request demonstrates action enforcement, not that the model ignored hostile text. Report unintended actions over all eligible hostile-workspace cases and clean completion over all ordinary cases. Restricting actions and reviewing conflicts can slow coding work. Encoded instructions, mislabeled trusted files, or an alternate tool route can bypass the tested path. A failure calls for stopping the session, revoking its credentials, and discarding unreviewed changes under the incident recovery process. Rerun both case groups before restoring that workspace path.

13.2 Generated code review

Code written by an agent can compile cleanly and still weaken security. Generated code is a proposed software change whose effect depends on the repository, build configuration, dependencies, and deployment target. Review has to establish what changed, which package identities entered the build, what secrets or permissions are accessible to the code, and which tests cover the affected behavior. Fluency and successful compilation supply little evidence about security.

The NIST Secure Software Development Framework4 calls for verification of third-party components across their life cycle and for review and testing of executable code. Pearce et al.5 found security weaknesses among GitHub Copilot suggestions in their August 2021 technical-preview study, using 89 constructed scenarios and 1,689 generated programs. Their result does not describe later versions, all languages, or workflows with independent testing.

The organization records the generated diff, dependency lock changes, scanner results, tests, reviewer decision, and deployment scope. Package references resolve to approved repositories and fixed identities before installation. Synthetic credentials test for accidental secret disclosure. These checks treat model output like any untrusted change while adding evidence about its generation path. The reviewed change can then enter a browser or execution environment with limited authority.

Before the release gate accepts a change, the build service checks the exact diff, dependency identity, test results, secret scan, reviewer decision, and target environment. Its review record links each changed file and dependency to the affected behavior, the tests run, and any remaining gap. Synthetic secret cases test the disclosure paths they actually exercise. Independent review and isolated builds consume compute and release time. Incomplete tests, build substitution, or a dependency that changes behind an accepted label can bypass the record. Recovery then blocks deployment, reverts the change, and restores the prior lock and artifact before the full gate runs again. Behavior and environments outside the tests stay unchecked.

13.3 Browser session authority

Brave6 reported on August 20, 2025, that hidden text on a Reddit page could steer Perplexity’s Comet browser agent. It made the agent send out the user’s email address and a one-time password (OTP). Brave is a competing browser vendor reporting its own test. Brave reported that the mitigation remained incomplete in its retest. The agent acted through the user’s signed-in sessions, so the question is which of those actions page text should ever be able to start.

A browser agent uses authority available through authenticated sessions, open tabs, browser storage, downloads, uploads, and clipboard access. A browser origin identifies a web protection domain based on scheme, host, and port. It helps isolate sites, but it does not express the employee’s business intent to submit a form or transfer data.

RFC 64547 defines origins as protection domains used by the same-origin policy. RFC 64548 also documents limits such as ambient authority and dependence on Domain Name System and transport security. Those browser facts do not form a complete authorization model for an AI agent using an already authenticated session.

In the example organization, the browser agent records the active tab, origin, session owner, target URL, form fields, upload or download path, and resulting state. Navigation and reading may fit one permission set. Submission, payment, file transfer, and clipboard use require separate checks. The cited browser and injection sources do not establish that these checks cover every action path in a particular browser agent. The same separation is needed when instructions arrive through images, audio, or video.

Each operation passes through the browser controller and action service. They check session identity, active origin, target origin, data movement, resolved object, and user approval. The test matrix counts blocked consequential actions over all forbidden cases, successful read-only tasks over their eligible set, and completed permitted consequential actions over all eligible permitted consequential cases. The permitted submission, payment and transfer cases check the destination, approved fields and resulting external state. Isolation and confirmation add friction and can break sites that change flows. Pop-ups, redirects, shared browser state, or an unmodeled clipboard path can bypass incomplete rules. Recovery closes the session, revokes its tokens, and withdraws transferred data where possible before the origin and action matrix runs again. Even a clean matrix does not show that page content was trustworthy.

13.4 Multimodal prompt injection

A multimodal model can receive text, images, audio, or video in one task. An application can also extract text from media before calling a text model. Both arrangements let external content affect the task, but they expose different intermediate records. The security decision concerns the source and authority of that content.

NIST’s adversarial machine learning taxonomy9 covers attack paths for multimodal models across several input types. OWASP’s 2025 prompt-injection entry10 discusses instructions hidden in images accompanying otherwise benign text. These references describe possible paths without giving a success rate for action-taking systems.

In an explicit extraction path, an optical character recognition processor converts image text into a text record, or a speech recognizer produces a transcript from audio. The application carries the media’s source and trust labels into that record before assembling model input. Retaining the original media, processor version, and any recorded page, region, or time offsets connects the text to what was processed. This path makes the conversion inspectable, but it can omit layout, sounds, or visual details needed to interpret the task.

A direct media path submits the image, audio, or video to a model that accepts that format. For example, Google’s Gemini documentation11 shows a request containing image data and a text question, without a separate application OCR step. The interface does not supply an intermediate transcript that the application can assume it has inspected. Evidence for this path therefore records the submitted media, any application resizing, cropping, or frame selection, accompanying text, model version, observed response, proposed action, and enforcement result. Internal processing that the service does not expose remains unknown. A text-extraction test cannot establish performance on this different input path.

Example

Tracing an image through two input paths. Suppose a support screenshot contains an error message and a low-contrast instruction to upload credentials. In an extraction test, the OCR processor returns both pieces of text. The application labels the result as untrusted screenshot content and retains it beside the original image. In a direct-input test, the application submits the image with the employee’s request and records the returned response. It has no separate OCR record unless a component actually produced one. In either test, any upload request still needs authorization at the action gateway. These are constructed test arrangements. Whether either model follows the hostile instruction must be observed.

For either path, the input handler preserves source labels and the action gateway checks the current task, proposal, and destination. The test corpus records the selected image, audio, and video cases, with clean controls for each supported path and form. None of the cited sources measures defense performance across all these modalities. Untested formats and transformations remain outside the result. Extraction and human inspection add compute and delay. Direct input leaves fewer intermediate records for diagnosis. Hidden content, parser differences, or a newly supported modality can fall outside the corpus. After a failure, the media is quarantined and dependent actions stop. Before the case runs through the full path again, the application owner restores affected application or session state where supported and records unresolved changes. The owner checks any completed external effect with its receiving service before arranging any needed authorized reversal or compensating action, where supported. A denied action does not prove correct interpretation of the media, while confinement can limit the effect of a mistaken interpretation.

13.5 Sandbox boundaries and resource limits

Hugging Face12 reported that an agent running an internal OpenAI evaluation escaped OpenAI’s evaluation sandbox and intruded into Hugging Face systems between July 9 and 13, 2026. Its reconstruction covers about 17,600 recovered actions and reports access to five customer datasets and operational search-query metadata. This is the affected company’s account, not an independent investigation. Escaping one boundary expanded the agent’s reach. Restrictions at later services still mattered.

An execution environment is the operating-system and network context in which agent-generated commands run. Its filesystem mounts, process rights, credentials, package installation, network routes, and time or resource limits define the intended execution boundary. If their enforcement fails, a command may reach beyond that boundary. The environment must match the task rather than the full authority of its host.

A host contains an isolated agent workload with a writable workspace, read only reference mount, credential scoped to inference, network allowlist, and CPU, memory, process, and time limits. The network path allows an inference service and blocks production and administration routes. Host secrets are shown outside the workload. File, process, network, and limit events feed an evidence recorder outside the workload. A separate recovery controller points to termination, credential rotation, and a replacement environment. No connection is drawn between the recorder and recovery controller, and no sockets are shown.
Figure 13.2: The figure’s ‘13.5 Execution environments’ heading refers to ‘Sandbox boundaries and resource limits’. An isolated execution environment bounds an agent process with specific mounts, credentials, network destinations, process rights, and resource limits. File, process, network, and limit events feed an evidence recorder outside the workload. A separate recovery controller can terminate the workload, rotate its credential, and start a replacement environment. The diagram shows no connection between recorder and recovery controller. Socket access restrictions require a separate check. Sockets are not depicted. The configured boundary depends on enforcement and does not establish resistance to every escape path.

NIST container guidance13 separates risks in images, registries, orchestrators, containers, and hosts, including shared-kernel and unrestricted-network concerns. NIST container guidance14 recommends controls at each layer and discusses separating workloads by sensitivity. Saltzer and Schroeder’s15 least-privilege principle supports granting only the rights needed for the task. These sources leave the escape resistance of a selected sandbox unproved.

Luo et al.16 showed that an agent can exhaust resources whose lifetimes span different scopes. A limit that resets after one turn does not bound what accumulates across a session or for the underlying process. The study used source-code-guided fuzzing to generate task-specific prompts for selected open-source web applications that use agents. It did not assume that the attacker had compromised the runtime. As a published demonstration on those applications it does not establish how often such exhaustion occurs elsewhere.

For a code repair in the example organization, the environment mounts one repository copy and supplies a short-lived test credential. It blocks production networks and caps processes, time, and storage. A container shares a kernel with its host, while a virtual machine adds a different isolation boundary at extra cost. None of the cited sources measures how often agents escape sandboxes or expose credentials. The bounded environment supports controlled multi-step action sequences. Enforcement must cover per-turn, per-session, and per-process budgets and account for shared resources that one task can leave for the next.

Example

Concrete confined-execution trace. The operating mode is an agent repairing a test repository. The task is to change one parser and run its unit tests. The reference environment allows writes under /workspace, access to packages.internal, and use of a test-only credential. It blocks production services and host storage.

Table 13.1: Agent execution trace.
Step Intermediate record Observed result
1 mount /workspace read-write parser file changed
2 mount /host absent read attempt returns path not found
3 destination packages.internal dependency request allowed
4 destination payments.internal connection denied
5 process count reaches 32 next process denied
6 unit test command 18 tests pass, 1 test fails

The agent reports a partial result and does not publish the change because the success condition required all 19 tests to pass. The conclusion is that the tested mounts, route rule, process limit, and stopping rule constrained this run. It does not exclude unknown escape paths or establish that the code change is secure. The paths, counts, and results are invented for the example.

Enforcement is spread over three places. The runtime checks mounts and process rights, the network gateway checks credentials and destinations, and the job controller applies resource limits. The confinement matrix records a defined set of permitted task operations and selected forbidden operations or escape attempts. Stronger isolation adds startup time, memory, and administrative work, and it can limit required tools. Kernel flaws, device access, or an unmetered route can bypass the tested boundary. Recovery terminates the environment and preserves its events. Exposed credentials are rotated, and the affected image or host is replaced before the matrix runs again. A clean matrix covers only the escape paths it actually tried.

13.6 Action sequence verification

In the Replit case at the start of this chapter, the delete happened during a code freeze. By Lemkin’s account, the agent’s later report that rollback would fail was wrong, so its claim about the outcome differed from the actual state.

A multi-step task can fail even when each command is allowed in isolation. An early action can change the state on which later authorization depends. A retry can repeat a non-repeatable operation, while a tool can report success before the external system settles. The sequence needs expected states, success checks, stopping rules, and a recovery path.

Complete mediation calls for checking each access instead of assuming that an earlier decision covers later access (Saltzer and Schroeder17). MCP18 recommends confirmation for tool invocations and preserving a person’s ability to deny an invocation. These principles do not define transaction handling, state drift, safe retry, or rollback for an agent workflow.

Debenedetti et al.19 built AgentDojo as a changing test environment for evaluating prompt-injection attacks and defenses for language-model agents. The benchmark checks legitimate task completion against attacker goals using mutable environment state. A separate check of resulting state is needed because a model that claims success may not have produced the intended effect. Chapter 14 (Evidence for release decisions) explains the research and defines the measurements for such evaluation. This chapter uses the same distinction between claimed and observed outcomes when verifying a sequence.

In the example, the coding agent reads the current repository state and proposes a patch. It applies the patch in the confined copy, then runs tests. After comparing the resulting diff, it stops before deployment. Each transition records its input state and observed result. An ambiguous result blocks automatic retry until the service can establish whether the action occurred. An idempotency key helps only when the receiver enforces it for a stable operation and payload, retains the key long enough, and defines failure semantics. It does not provide a universal exactly-once guarantee or an automatic undo. This example and the cited benchmark do not establish a general transaction guarantee for agent workflows. Chapter 14 treats this full sequence as the system behavior that must pass evaluation before release.

The workflow controller checks each transition against the prior verified state, current authorization, action budget, expected result, and retry key. Its tests count correctly stopped or recovered failures over all injected sequence failures, and successful completion over clean workflows. Verification adds service calls and can leave work unfinished when external state settles slowly. Eventual consistency, missing result fields, or an irreversible side effect can defeat automatic recovery. When possible, recovery resumes from the last verified state. Otherwise it applies a supported compensating action or hands the case to an authorized person. A submitted call, an acknowledged operation, a completed effect, and settlement are four distinct checks in the record. One safe sequence does not show that other tools combine safely.

Note

Chapter checkpoint. In the example organization, a coding agent must repair a parser and install an approved dependency. It must then run tests and return a patch. The repository also contains issue text instructing the agent to upload credentials. What execution and verification record supports a bounded result?

Answer. Keep the employee’s repair request as the trusted task and record the issue text as untrusted workspace data. Mount only the repository copy, provide a short-lived test credential, permit only the approved package destination, block production routes, and set process, time, and storage limits. Record each proposed command, policy decision, filesystem change, network result, and test outcome. Return the patch as a completed repair only if the stated tests pass. If the agent attempts an upload, retain the submitted request and observed denial as evidence of enforcement. The presence of hostile issue text alone does not establish that an upload was attempted. This supports the observed run and does not establish that the sandbox has no unknown escape path.

The AgentDoS demonstration and the AgentDojo benchmark add specific evidence on resource budgets and on checking claimed versus observed task outcomes. These sources do not establish protection across all coding workspaces, browser agents, or multimodal inputs. They also do not supply a general sandbox escape rate or a transaction guarantee for arbitrary multi-step workflows.


  1. Beatrice Nolan (2025), “AI-powered coding tool wiped out a software company’s database in ‘catastrophic failure’,” Fortune, July 23, 2025, source. News coverage of a first-person account, with no forensic report behind it.↩︎

  2. Kai Greshake et al. (2023), “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec), 79-90, DOI, author manuscript arXiv:2302.12173v2, author-manuscript section 3, PDF pp. 3-5, source.↩︎

  3. OWASP GenAI Security Project (2025), Top 10 for LLM Applications 2025, LLM01:2025, “Prompt Injection,” “Indirect Prompt Injections,” and “Example Attack Scenarios,” source.↩︎

  4. Murugiah Souppaya, Karen Scarfone, and Donna Dodson (2022), Secure Software Development Framework (SSDF) Version 1.1: Recommendations for Mitigating the Risk of Software Vulnerabilities, NIST SP 800-218, practices PW.4.4 and PW.7.1-PW.8.2, printed pp. 13-15, source.↩︎

  5. Hammond Pearce, Baleegh Ahmad, Benjamin Tan, Brendan Dolan-Gavitt, and Ramesh Karri (2022), “Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions,” 2022 IEEE Symposium on Security and Privacy, 754-768, DOI, author manuscript arXiv:2108.09293, author-manuscript sections III-V, PDF pp. 3-13, and section VI, “Threats to Validity,” PDF pp. 13-14, author manuscript.↩︎

  6. Artem Chaikin and Shivan Kaul Sahib (2025), “Agentic Browser Security: Indirect Prompt Injection in Perplexity Comet,” Brave Blog, August 20, 2025, source. Written by a company that sells a competing browser, about one attack path and its own retest.↩︎

  7. Adam Barth (2011), The Web Origin Concept, RFC 6454, IETF, sections 3.1-3.4, pp. 4-8, source.↩︎

  8. Adam Barth (2011), The Web Origin Concept, RFC 6454, IETF, sections 8.1-8.3, pp. 14-16, source.↩︎

  9. Apostol Vassilev et al. (2025), Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, section 4.2.3, printed pp. 58-59, source.↩︎

  10. OWASP GenAI Security Project (2025), Top 10 for LLM Applications 2025, LLM01:2025, “Prompt Injection,” multimodal discussion and “Scenario #7: Multimodal Injection,” source.↩︎

  11. Google, “Image understanding,” Gemini API documentation, “Passing images to Gemini” and “Passing inline image data,” retrieved September 24, 2026, source. Documents an image-plus-text API input, not the service’s complete internal representation or resistance to prompt injection. The evidence-record comparison is an engineering consequence of which intermediate values the application actually receives.↩︎

  12. Hugging Face (2026), “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident,” Hugging Face Blog, July 27, 2026, source. The affected company’s own reconstruction of the incident, not an independent investigation. OpenAI, operator account, July 21, 2026, “What happened during this incident” and July 28 update, operator account. These are first-party accounts. Recovered actions are not a count of successful exploits.↩︎

  13. Murugiah Souppaya, John Morello, and Karen Scarfone (2017), Application Container Security Guide, NIST SP 800-190, sections 3.1-3.5, especially 3.4.2 and 3.5.2, source.↩︎

  14. Murugiah Souppaya, John Morello, and Karen Scarfone (2017), Application Container Security Guide, NIST SP 800-190, sections 4.1-4.5, especially 4.3.4 and 4.4.2, printed pp. 23-24, source.↩︎

  15. Jerome H. Saltzer and Michael D. Schroeder (1975), “The Protection of Information in Computer Systems,” Proceedings of the IEEE 63(9), 1278-1308, section I.A.3, “Design Principles,” item f, “Least privilege,” source.↩︎

  16. Jiaqi Luo et al. (2026), “Autonomy Comes with Costs: Detecting Denial-of-Service Vulnerabilities Caused by Resource Abusing in LLM-based Agents,” Proceedings of the USENIX Security Symposium 2026, 3991–4010, sections 2.1–2.4, 4.1.1, 6.1 and 9, official publication and paper. Limit: studied selected open-source agent web applications with source-code-guided fuzzing and did not estimate prevalence for arbitrary external model services or compromised runtimes.↩︎

  17. Jerome H. Saltzer and Michael D. Schroeder (1975), “The Protection of Information in Computer Systems,” Proceedings of the IEEE 63(9), 1278-1308, section I.A.3, “Design Principles,” item c, “Complete mediation,” source.↩︎

  18. Model Context Protocol contributors (2026), Model Context Protocol Specification, revision 2026-07-28, “Tools,” “User Interaction Model,” source.↩︎

  19. Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer and Florian Tramèr (2024), “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents,” Proceedings of the Thirty-eighth Conference on Neural Information Processing Systems, Datasets and Benchmarks Track, sections 3, 3.1, 3.4 and Figure 1, proceedings paper. Limit: evaluation environment with specified tasks and attacker goals, not a deployment incident rate. Full method and measurement definitions are in Chapter 14.↩︎