9  Runtime intrusion and resource abuse

Exposed interfaces, stolen credentials, unrestricted network access, and excessive work can compromise an AI service or make it unavailable to legitimate users.

Oligo Security1 reported on March 26, 2024, that attackers were exploiting exposed clusters of Ray, an open-source framework that spreads AI jobs across many machines. Ray’s Jobs API accepted job submissions without authentication, and Oligo traced the first attacker foothold it found to September 5, 2023. Oligo estimated the value of machines and compute that might have been compromised at almost $1 billion. The behavior is tracked as CVE-2023-48022, but Ray’s developers dispute that it is a flaw. They say the behavior is expected because Ray is meant to run on a trusted network. Whichever side is right about the label, an interface that accepts jobs from anyone who can reach it gives those callers code execution and compute.

An AI job service accepts requests, gives workers access to data, and allocates compute. An attacker may try to enter through an exposed administration interface, reuse a stolen worker credential, send data through an allowed network path, or consume the capacity needed by other jobs. These routes affect different controls. Restricting one endpoint does not, for example, limit the compute consumed by a caller who is already authorized.

Assume the organization operates the job service and model server on rented infrastructure. Each worker runs with a service identity. The execution-isolation controls (Chapter 8) limit what its process can reach on that platform. Service security adds the checks on incoming calls, credential use, network destinations, and resource consumption, together with a way to stop harmful jobs while retaining essential service.

An evaluation job passes endpoint and credential checks into a confined worker with approved storage and model routes. Network, host, privilege, and resource controls block separate paths. An incident trace records their events and distinguishes exposure, execution, effects, and missing evidence.
Figure 9.1: The ‘Containing runtime exposure’ overview introduces Chapter 9, ‘Runtime intrusion and resource abuse’. Runtime containment follows a legitimate job through reachable interfaces, workload identity, permitted network flows, process isolation, and scheduler limits. Each enforcement point produces evidence for a common incident trace. Reachability, execution, and downstream effects remain different observations, and missing telemetry remains an explicit gap. These records support containment and rebuilding under the tested configuration.

9.1 Exposed job and admin interfaces

ShadowRay began at a job interface that nobody had placed behind an access check. Other AI services show the same gap. Langflow, a tool for building AI workflows, had an unauthenticated endpoint that ran submitted code, tracked as CVE-2025-3248 and published on April 7, 2025 (NIST National Vulnerability Database2). The US Cybersecurity and Infrastructure Security Agency (CISA) added it to its Known Exploited Vulnerabilities catalog on May 5, 2025. Trend Micro3 reported on June 17, 2025, that attackers were using it to install the Flodrix botnet. Separately, Wiz4 reported on January 29, 2025, that DeepSeek had left ClickHouse databases reachable without authentication. That case exposed stored data rather than code execution, but the same kind of overlooked interface lacked an access check.

An overlooked interface can let calls skip checks on better-known endpoints because the inventory missed that interface. A deployment interface accepts commands or configuration for workloads. An administrative endpoint can change platform state. A job-submission API accepts work and parameters for execution.

NIST container guidance5 treats orchestrators and runtime interfaces as high-value control surfaces and identifies unbounded administrative access as an orchestrator risk. Kubernetes guidance6 states that kubelet API access should be restricted and not exposed publicly, a checklist item the team must apply to the deployed Kubernetes version. Neither source inventories the AI notebook, scheduler, or model-server products in a given deployment, so product-specific checks remain necessary.

Reviewers enumerate notebook servers, model endpoints, schedulers, orchestration interfaces, management consoles, and health or metrics endpoints, recording network reachability and permitted operations for each. The design under review places a gateway and network policy in front of them, configured to admit a call only for a known caller identity from a permitted source network requesting an allowed operation on the intended endpoint. A receiving service should reject calls that fail those checks.

Example

Reviewing deployment interfaces. Suppose reviewers compare a job service inventory with reachability tests from an external network and an approved internal network. They also test the model and job APIs using controlled identities. The model API is available through an internal gateway, and the job API through the internal network. An administration page responds from the internet even though its inventory entry permits only internal access. No administration login is attempted.

Table 9.1: Deployment interface comparison.
Interface record Intended caller Observed reachability Authentication result Documented capability
model-api support application internal gateway only employee session rejected Not exercised by the rejected caller
job-api build service internal network only build identity accepted submit job
admin-ui platform operator internet reachable Not tested; login page returned Change configuration; not attempted

Conclusion. The admin-ui route is an exposure that needs removal or a justified access path. It is not evidence that anyone authenticated or changed state through it.

The result is unexpected reachability among the inventoried interfaces, recorded separately for each tested network. Finding zero unexpected routes after a repair means none was found within that inventory and those test locations. Intended routes may remain available. Restricting administration can slow urgent maintenance, while an omitted endpoint or an alternate route remains outside this result.

9.2 Worker execution boundaries

After a call passes the gateway, the worker process may need files, devices, and network access to perform the job. Its service identity supplies a calling identity for receiving services, which independently check its credentials and requested operations. Those checks do not restrict the worker’s host mounts or process privileges. The execution boundary applies those limits. Authorization and execution limits remain separate checks.

9.3 Workload identities and credentials

Hugging Face7 disclosed on May 31, 2024, that secrets stored for its Spaces, applications hosted on its platform, could have been accessed without authorization. The company revoked tokens in response. The disclosure did not identify the attacker, describe the method, or say how many secrets were affected. A stolen bearer credential can let its holder use the workload’s access while receiving services still accept it. The result depends on the credential’s scope, expiry, revocation handling, and any additional caller checks.

Shared or long-lived credentials make it difficult to limit or attribute what a service does. A workload identity identifies a non-human process or service. A service account is one way to give a workload such an identity. Token lifetime is the period for which a credential can authorize requests.

Kubernetes service accounts supply workload identities that can be bound to permissions through role-based access control (Kubernetes documentation8). NIST’s cloud-native zero-trust model9 recommends short-lived, cryptographically verifiable service identities with service-to-service authorization. A service account remains an identity mechanism only: secure permissions, token handling, and any external federation stay deployment choices, and the recommendations do not prove a particular identity plane cannot be bypassed.

The design associates each interface call with a workload identity and checks its role before the operation runs. Short-lived credentials arrive through a controlled mount or exchange. Their audience identifies the service allowed to accept them, and their scopes describe allowed kinds of access. Both should be limited to what the job needs.

Example

Tracing a job token. Suppose an employee submits a request to evaluate a test set. The application creates job-417 and obtains a worker token with subject eval-worker, audience job-api, scope jobs:read,results:write, and an expiry fifteen minutes after issue. The job service verifies the issuer and signature, checks that the token is current and intended for that service, and checks the requested operation against both the scopes and the worker’s assignment to job-417. Assume those checks pass for the assigned read.

The worker then presents the same token to model-admin. That service rejects it because the audience differs and the token provides no administration permission. The example tests two service decisions. It does not show that the credential cannot be stolen or that every service validates it correctly.

The receiving service should validate the credential, its issuer, intended audience, lifetime, and the requested resource and operation. If the design supports revocation, the service also needs a way to obtain and enforce that state. A self-contained token checked only offline may remain accepted until expiry unless an additional deny rule is applied. Short lifetimes reduce that window but add issuance traffic and dependency on the identity service. Tests therefore cover the actual service-operation pairs and credential states, including expired or revoked credentials where those mechanisms exist. Disabling future access does not erase data already copied by a worker.

9.4 Network flow controls

A workload identity loses much of its value if traffic can simply bypass the enforcement point. Ingress enters a service boundary, while egress leaves it. Control traffic changes system state or configuration. Data traffic carries prompts, model inputs, outputs, or stored records.

NIST SP 800-207A10 places cloud-native access policy around application and service identities, with gateways or proxies enforcing identity and network policy. This reference architecture does not prove a deployed proxy cannot be bypassed. NIST container guidance11 recommends controlling container egress with application-aware filtering where changing addresses limit traditional network rules, without supplying one policy for every AI data and control flow.

The design maps each required flow to a source identity, destination, protocol, and purpose, denies flows outside that map, and records each decision for review.

Example

Tracing evaluation traffic. Suppose a worker must read a test object, call an internal model endpoint, and write a result, with the approved flow record naming those destinations and denying other egress.

Table 9.2: Network exposure event trace.
Event Source identity Destination Policy result Observed result
1 eval-worker test-store allow read object case-18 returned
2 eval-worker model-api allow request response resp-903 returned
3 eval-worker results-store allow write result run-417 stored
4 eval-worker 203.0.113.25 deny connection blocked

Conclusion. The blocked fourth connection demonstrates enforcement on the tested path. It does not explain why the process attempted that connection, nor whether alternate routes are closed.

A test can report prohibited connections blocked out of all attempted prohibited connections, alongside required connections completed out of all attempted required connections. Mixing both groups into one blocked-traffic fraction would obscure whether either rule worked. Narrow rules need maintenance as destinations change. DNS changes, an unchecked proxy, or another network interface may create a path outside the tested policy.

9.5 Runtime updates and resource limits

Oligo Security12 reported on November 18, 2025, a second campaign against exposed Ray clusters, which it called ShadowRay 2.0. The attackers, tracked by Oligo as IronErn440, took over exposed clusters, and the campaign spread itself from compromised clusters to further ones. They hosted their code first on GitLab and then on GitHub. Oligo counted more than 230,000 Ray servers exposed to the internet. Oligo judged it likely that the attackers used AI to write their code, a point taken up in Chapter 17 (AI-assisted attacks). Oligo said the campaign could have been active since September 2024. Hijacked compute and a runaway legitimate job show the same symptom: work consuming capacity that nobody approved.

A legitimate job can consume excessive capacity or cost, even when its credentials and software are approved. A long generation loop, many concurrent jobs, or repeated tool calls can prevent other work from finishing. A runtime patch updates execution software to correct a defect. A resource quota limits consumable compute, memory, storage, or concurrent work. Runaway computation continues beyond the intended cost or time.

NIST guidance13 recommends monitoring container runtimes for vulnerabilities, remediating findings, and blocking deployment to unmaintained runtimes. It does not set a patch deadline or establish effectiveness in a particular cluster. NCSC guidance14 recommends monitoring behavior and treating changes to data, models, prompts, or software as updates that need testing and version handling. It does not define accelerator quotas or a scheduler-specific containment method.

The working design pairs an approved runtime version and vulnerability response with job timeouts, quotas, concurrency limits, and billing alerts. The scheduler admits further work or stops running work based on job identity, requested resources, current use, elapsed time, and concurrency state. Accelerator and scheduler limits beyond that need the product documentation and direct tests for the selected platform.

Example

Testing time and concurrency limits. Suppose a batch inference job is configured to request repeated generation without a stopping condition, against reference limits of ten minutes, eight gigabytes of memory, and two concurrent workers. At minute six the job holds 7.4 GB. At minute eight its request for a third worker is denied. At minute ten the scheduler stops the remaining processes and records limit=time.

Conclusion. The configured time and concurrency controls produced a bounded failed job instead of completion. The numbers test this configuration only and do not establish suitable limits for a real model or accelerator.

The test reports stopped over-limit jobs out of all intentionally over-limit jobs, alongside clean completions out of all ordinary cases. A justified limit change requires both groups to run again under the revised threshold. The result still does not transfer to a different workload. Tight limits reduce cost exposure but can reject legitimate long work, while an unmetered service or work split across identities escapes a local quota.

9.6 Runtime containment and continuity

An exposed endpoint shows reachable capability, not everything that happened afterwards. Execution exposure is an interface condition that permits unauthorized or unsafe execution. Propagation is movement from the first affected workload to another system. Attribution evidence supports a conclusion about who caused an event and with what confidence.

NCSC guidance15 recommends logging inputs and monitoring behavior to support audit, investigation, and remediation. NIST container guidance16 identifies misuse of a container to attack other containers, hosts, or systems as a possible failure. Establishing that such propagation or resource theft occurred requires evidence from the affected system.

Investigation orders the exposed route, authentication records, request logs, workload creation, process events, credential use, network flows, and resource consumption by time. Missing telemetry marks a gap in the record rather than proof nothing happened there. Runtime or tool output from outside the trust boundary keeps its source label when it enters the data path, and the investigation record retains that label.

Example

Replaying a resource test after containment. Suppose responders stop the over-limit batch job and approve a shorter limit of eight minutes. They preserve the available worker events, cancel related harmful jobs, and compare recorded usage with billing records. A test repeats both the intentionally over-limit job and ordinary jobs under the new setting.

Assume the over-limit job is stopped at eight minutes and all ordinary test jobs finish within six minutes. Those observations support the new limit for these cases. They do not establish who caused the original runaway or exclude activity outside the recorded window.

Stopping one workload can release capacity for essential services, but stopping a shared identity service or network path may interrupt those services too. A continuity plan can specify which operations remain available and which may be suspended. Containment and evidence collection need a case-specific order: responders preserve volatile records where feasible without prolonging active harm, but an ongoing disclosure may require blocking access first. Revocation does not inherently erase logs, and preserving evidence is not a universal reason to delay it.

After restricting the affected paths, responders can rebuild a worker and repeat the relevant access and resource tests. An unknown credential or an unlogged route may remain outside that verification. Restoring the worker cannot undo earlier transmissions or business operations. A host operator can also delay or stop the workload, so application-level continuity remains dependent on the underlying platform. The fuller response method is developed in incident recovery.

The service record supports task and action budgets in Chapter 12 and execution limits in Chapter 13. The same record supplies conditions for release tests, observations for live monitoring, and evidence for incident recovery.

Note

Chapter checkpoint. A review finds an internet-reachable administration page, a shared worker credential, unrestricted egress, a host filesystem mount, and no job timeout. Which records and changes would show the runtime contained for an evaluation job?

Answer. Record the intended caller and permitted operation for every interface, then remove the public route or place it behind an approved access path. Give the evaluation worker a short-lived identity limited to the job and result services. Allow only the required storage and model flows, remove the host mount, and set tested time, memory, and concurrency limits. Repeat the interface, identity, network, isolation, and resource traces and retain each allow or deny result. Passing those checks supports only the tested configuration. It does not prove the runtime has no unknown vulnerability or alternate route.


  1. Oligo Security, “ShadowRay: First Known Attack Campaign Targeting AI Workloads Exploited In The Wild,” Oligo blog, March 26, 2024, source. The almost $1 billion figure is Oligo’s estimate of the value of machines and compute that might have been compromised, and Ray’s developers dispute that CVE-2023-48022 is a vulnerability.↩︎

  2. National Institute of Standards and Technology, “CVE-2025-3248 Detail,” National Vulnerability Database, published April 7, 2025, source. The entry records the flaw and its catalog status, not the extent of real attacks.↩︎

  3. Trend Micro Research, “Critical Langflow Vulnerability (CVE-2025-3248) Actively Exploited to Deliver Flodrix Botnet,” Trend Micro, June 17, 2025, source. This is one security vendor’s observation of one campaign, not a count of affected Langflow servers.↩︎

  4. Wiz Research, “Wiz Research Uncovers Exposed DeepSeek Database Leaking Sensitive Information, Including Chat History,” Wiz blog, January 29, 2025, source. The report describes open databases found by researchers, not code execution or confirmed misuse by attackers.↩︎

  5. Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Section 3.3, especially 3.3.1, printed pp. 15-16, source.↩︎

  6. Kubernetes contributors, “Security Checklist,” Kubernetes documentation (living page), “Network security” and “Authentication and authorization,” source. This living checklist must be applied to the deployed Kubernetes version.↩︎

  7. Hugging Face, “Space secrets leak disclosure,” Hugging Face blog, May 31, 2024, source. The disclosure does not identify the attacker, the method, or the number of secrets accessed.↩︎

  8. Kubernetes contributors, “Service Accounts,” Kubernetes documentation (living page), “What are service accounts?” and “Grant permissions to a ServiceAccount,” source.↩︎

  9. Ramaswamy Chandramouli and Zack Butcher, A Zero Trust Architecture Model for Access Control in Cloud-Native Applications in Multi-Location Environments, NIST SP 800-207A (2023), Section 3.1, ID-SEG-REC-2 and ID-SEG-REC-3, printed p. 8, source.↩︎

  10. Ramaswamy Chandramouli and Zack Butcher, A Zero Trust Architecture Model for Access Control in Cloud-Native Applications in Multi-Location Environments, NIST SP 800-207A (2023), abstract and Sections 2-3, source.↩︎

  11. Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Section 4.4.2, printed pp. 24-25, source.↩︎

  12. Oligo Security, “ShadowRay 2.0: Active Global Campaign Hijacks Ray AI Infrastructure Into Self-Propagating Botnet,” Oligo blog, November 18, 2025, source. The server count measures exposure, Oligo says AI-generated code is only strongly implied, and the September 2024 start date is only possible.↩︎

  13. Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Section 4.4.1, printed p. 24, source.↩︎

  14. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), Secure operation and maintenance, PDF p. 16, source.↩︎

  15. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), “Monitor your system’s behaviour” and “Monitor your system’s inputs,” PDF p. 16, source.↩︎

  16. Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Section 3, printed p. 13, source.↩︎