8  Isolation on shared infrastructure

Data processed by an AI service can be exposed or altered through the platform on which it runs.

An approved model can still process private data on a platform that exposes them. Another customer’s process may reach across a workload boundary. A privileged host component may inspect memory the workload cannot hide, or state left behind in shared memory may become readable to whoever runs next. Controls between tenants, the customers sharing the platform, limit what another workload can reach. Separate checks determine who can decrypt data and what evidence supports the platform’s claimed state. Readable data may still remain in the application or its logs. Tenant isolation, decryption controls, and evidence about platform state answer different questions, so a passing result from one cannot substitute for the others.

Assume the organization operates its model workload on rented infrastructure. The provider operates the hardware and firmware beneath that workload, and the deployment record states who operates each layer before runtime access is configured. The organization can inspect only its own workload records and whatever platform evidence the provider supplies, not the provider’s administrators, firmware internals, or other tenants. A provider-operated model API is the contrasting arrangement used for comparison below. There the provider operates the model service and its platform, while the organization may still operate the calling application, user identities and credentials, context assembly, and transfer policy.

Three deployment stacks assign customer and provider duties. An expanded workload shows tenant boundaries, encrypted storage, and protected execution. An attester supplies evidence to a verifier, which sends a result to the key service. Audit scope marks unresolved side channels, availability, outputs, and future state. Correctly numbered panels include GPU local-memory reuse and memory corruption, with affected-stack and access conditions. A key-service-to-workload arrow shows the explicitly qualified raw-key export variant.
Figure 8.1: The ‘Limiting platform access’ overview introduces Chapter 8, ‘Isolation on shared infrastructure’. Deployment choices move operating duties between the organization and provider. Tenant separation, encryption, and protected execution address different access paths. Remote attestation carries measured-state evidence to a verifier, whose result informs a separate key-release decision. That result covers the appraised claims and accepted trust assumptions, leaving later state changes and excluded channels unresolved. The pictured raw-key export is an alternative to the K-18 example, which keeps the key in the service. Expiry cannot erase already released material. GPU exposure examples retain their specific hardware and attacker-access conditions.

8.1 Tenant isolation

Wiz researchers1 reported on April 4, 2024, on authorized research against the Hugging Face Inference API. They uploaded a malicious model in pickle format and used it to run their own code on the service. Chapter 7 (Supply chain compromise) explains how pickle files can run code when loaded. Wiz said that reaching cluster metadata in Amazon EKS, Amazon’s managed Kubernetes service, from that position could have allowed access to other customers’ workloads. Wiz described that cross-tenant access as possible, not as something it carried out.

On September 26, 2024, Wiz disclosed CVE-2024-0132 in the NVIDIA Container Toolkit, the software that gives containers access to NVIDIA GPUs. It is a time-of-check to time-of-use flaw, in which a condition checked at one moment changes before use, and it allowed escape from a container to the host. The flaw carries a Common Vulnerability Scoring System (CVSS) severity score of 9.0 out of 10 (Wiz report, NVIDIA advisory, and bypass disclosure2). A February 2025 disclosure by a researcher reported a bypass of the first fix, tracked as CVE-2025-23359. These disclosures identify particular escape paths. They do not establish how often attackers used them.

In both cases the separation between tenants depended on specific platform software, not on the label “container” or “managed service.” An evaluation job can run beside jobs belonging to strangers, and a submitted script may itself process hostile input. The first boundary to test is therefore the one between workloads sharing a platform.

Tenant isolation is that separation across compute, memory, storage, network, and management paths. A container normally isolates processes while sharing the host operating-system kernel. A virtual machine runs a separate guest operating system behind a hypervisor-managed boundary instead, with the hypervisor mediating CPU, memory, devices, storage, and networking. These labels identify different enforcing mechanisms. Neither proves that the deployed tenant boundary blocks every cross-workload path. The result for a virtualized boundary depends on the exact hypervisor, privileged management components, device-assignment mode, configuration, and patch state (NIST hypervisor guidance3).

NIST container guidance4 states that containers share a kernel and identifies risks around containers, hosts, images, registries, and orchestrators. NIST container guidance5 recommends layered isolation with application-aware network controls rather than simple host separation alone. NIST container guidance6 names insecure defaults, excessive privileges, host file access, runtime vulnerabilities, and shared-kernel exposure as container risk classes. It recommends separating workloads by sensitivity while constraining privileges, host access, and network access. A privilege is an allowed operation that exceeds ordinary process access, and a host mount exposes host storage inside a workload. Limits on privileges, host access, and network access reduce paths without establishing complete containment, especially with a shared kernel. The long-standing least-privilege principle behind them holds that each workload should receive only the rights its task needs (Saltzer and Schroeder7).

No single component draws the whole boundary. The hypervisor, the container runtime, the storage service, and the network policy each own part of it. Tenant identities gate cross-tenant access, and memory and storage objects carry ownership. Routing configuration determines which network paths are available. Network and service policies govern permitted access, and management permissions determine who may reconfigure these controls. Within a workload, execution isolation limits the processes, files, devices, memory, and network resources the workload can affect. One common arrangement runs the workload as a restricted user with limited capabilities, then narrows further through approved mounts, device access, network policy, and resource limits. Passthrough devices and management paths need their own evidence rather than inheriting the container or VM claim. Agent-generated commands run under these same controls when the task calls for executing model-produced code. The record of what that particular task was allowed belongs to the agent sequence discussion in agent environments (Chapter 13).

A tenant-separation trial seeds cross-tenant attempts that the policy should deny alongside permitted accesses it should allow. It reports blocked attempts over all seeded attempts next to successful permitted accesses over all controls. Stronger separation costs sharing: dedicated boundaries consume memory and startup time, reduce accelerator sharing, and can limit accelerator access. The trial stays bounded by what it attempted. A host flaw, residual device state of the kind examined below, or a management path outside the test can carry access past the tested rules. Passing says nothing about accelerator or provider properties the trial never exercised.

Example

Testing mount and privilege boundaries. Suppose reviewers execute a generated analysis script that reads the mounted repository, then reaches for a host secret and requests an added privilege. The reference policy allows reads under /workspace with ordinary processes only. The worker opens /workspace/input.csv successfully, the open of /host-secret returns access denied, and the privilege request is denied as well before the job exits after writing its ordinary result.

Conclusion. The tested mount and privilege controls held for these two attempts. The run says nothing about the shared kernel or runtime having no other escape path.

8.1.1 GPU local-memory reuse

Residual GPU data creates a different cross-workload confidentiality path. In the 2024 LeftoverLocals work8, uninitialized GPU local memory remained readable across invocations of GPU kernels on tested hardware and software stacks. A GPU kernel is a program that runs on the GPU. An attacker with access to a shared programmable GPU interface could observe data another kernel left in that region. CERT/CC9 tracks the issue as VU#446598 and CVE-2023-4969 with vendor-specific status, not as a finding about every GPU or memory region. The claim therefore specifies the local-memory region, the kernel-to-kernel reuse event, the attacker’s GPU API access, and the tested GPU, operating system, driver, and firmware. Volatile local-memory clearing and storage-media sanitization are separate controls.

Clearing is a control to verify. The evidence must identify the exact mitigation, whether it is enabled by default, the event that triggers it, its order relative to the next workload, and how completion is observed. AMD’s bulletin10 describes a mode, disabled by default, that an administrator can enable on supported products. It is designed to prevent concurrent GPU processes and clear registers between processes, with possible performance effects. That statement supplies conditions for those supported products. It does not establish a universal register-clearing property.

Example

Testing GPU local-memory reuse. Suppose reviewers specify the GPU, driver, runtime, and scheduling configuration for a controlled test. In this setup, a first tenant’s kernel writes a known pattern into a specified local-memory region and reads it back to confirm that the test can detect the pattern. The scheduler then runs a second tenant’s listener kernel immediately afterward on the same tested compute path, using a recorded order and no unrecorded intervening work. The policy forbids the second tenant from reading any byte of the first tenant’s pattern.

If the listener detects the pattern, the test has observed a cross-tenant confidentiality failure. If it does not, the result is only non-detection for the tested region, order, device, driver, runtime, and scan coverage. It does not prove that clearing occurred or that other regions and schedules are safe.

8.2 Provider administrator access

Ordinary container and virtual-machine isolation relies on privileged platform components. Administrators who can change hardware, firmware, virtualization, storage, networks, or identity policy may have access beyond the tenant’s process-level rules. Confidential-computing designs try to reduce some of those paths under specific assumptions.

Shared responsibility divides security duties between customer and provider. NCSC guidance11 states that modern AI systems commonly combine software, data, models, and remote services from several parties, which makes security responsibility harder for users to inspect. NIST zero-trust guidance12 removes implicit trust based only on network location, affiliation, or ownership and focuses access decisions on resources, while still depending on administrators, hardware, identity systems, and policy enforcement.

The control plane manages resources, identities, and configuration, and administrative access is the authority to change or inspect those layers. The administrative threat depends on who can change the host, hypervisor, firmware, device assignment, storage, network, verifier policy, reference values, and key service. Ordinary containers and VMs leave host, hypervisor, device, storage, network, and management administrators inside the trust boundary. A confidential VM can reduce some normal administrative read paths, but only under the specific hardware and software threat model selected for that deployment. Rented infrastructure alone does not protect plaintext from provider administrators (Confidential Computing Consortium analysis13). Which duties the provider claims and which stay with the customer still comes from the provider’s own statements, not from the arrangement label (NCSC guidance14).

Example

Comparing three arrangements. Suppose reviewers place the same inference task behind an external model API, on rented infrastructure, and on premises, listing every layer from hardware through identity. With the external model API, the provider operates the model service and its platform. The organization operates the calling application, user identities and credentials, and transfer policy in this example. On rented infrastructure the organization also operates the guest system and model service. On premises it additionally operates hardware and firmware, although suppliers still provide components and updates.

Conclusion. Each arrangement still requires trust in some external components or parties. The comparison shows where each administrator path sits rather than eliminating any of them.

8.3 Key custody and decryption authorization

An encrypted document becomes readable somewhere when a model processes it. Reviewing who can obtain that plaintext requires both the data state and the key holder. Data at rest is stored, data in transit moves between endpoints, and data in use is being processed in memory. Key custody identifies which party or component can use or release a cryptographic key. NIST key-management guidance15 treats generation, distribution, storage, use, backup, recovery, revocation, and destruction as managed lifecycle functions, without establishing that a provider lacks access to plaintext or keys.

Before decryption, the service may also require evidence about the requesting workload’s state. An appraisal evaluates that evidence against a policy and approved reference values. The component performing this check is the verifier. Its result informs the key service’s separate access decision. It does not itself grant document access or prove safe future behavior (RATS architecture16). The complete evidence flow is explained in Remote attestation.

Suppose the design keeps document key K-18 inside a key service. Workload W-4 asks for a five-minute decryption session. Before opening it, the service checks current workload identity and current authorization for document D-18. It also checks the requested purpose against that authorization and requires an acceptable appraisal result. It checks authorization again for each decrypt request during the session and rejects new requests after the deadline. The gate checks each requested decryption against identity, authorization, purpose, appraisal, and deadline. It cannot enforce the later purpose of plaintext already returned. This design does not export raw key bytes. If another design does export them, expiry at the service cannot erase copied key material. NIST zero-trust guidance17 supports authentication and authorization before a session to an enterprise resource, but does not specify this confidential-computing protocol.

Example

Tracing a denied decryption session. Suppose inference runs on rented infrastructure and workload W-4 requests a five-minute session for document D-18. The key service confirms identity and current document authorization and checks the requested purpose against that authorization, but receives no acceptable appraisal result. It denies the session, records each checked input, and does not decrypt the document.

Conclusion. The gate denied this path while the document remained encrypted. It says nothing about an alternate decryption path, copied key material in a different design, or plaintext retained after a permitted decryption.

Short sessions and repeated appraisal add latency and availability risk. Operating more layers directly adds patching and on-call work. Relying on a provider for those layers leaves the organization with less direct evidence. The release check stays bounded by its inputs. An alternate decryption path, a copied session credential, plaintext retained in process or GPU memory, output, or logs, or an incorrectly bound identity can carry access beyond the gate.

8.4 GPU memory corruption

Even with tenants separated and keys controlled, hardware faults can corrupt a workload’s active state. Working memory here means the parameters, inputs, and intermediate values a workload holds while it runs.

GPUHammer18, presented at USENIX Security 2025, demonstrated an integrity attack on an NVIDIA A6000 with GDDR6 under the paper’s attacker access, co-location, and placement conditions. Rowhammer is a hardware fault mechanism in which repeatedly accessing DRAM rows can induce bit flips in nearby rows. The study produced such bit flips and showed model damage on its tested setup. It does not establish that every GPU is vulnerable, and its reported accuracy change does not predict operational harm elsewhere. Error-correcting code (ECC) memory adds redundant information so that specified memory errors can be detected or corrected. The GPUHammer researchers19 identify memory-controller ECC as a mitigation for their A6000 attack. They report that ECC is disabled by default on that platform, consumes some usable memory when enabled, and can slow workloads. These conclusions apply to that platform and attack, not GPU Rowhammer as a class.

The platform record for this risk specifies the GPU and memory type, ECC mode, firmware and driver, attacker execution access, co-location condition, observed fault, and integrity check. A clean result on another device is evidence only for that tested setup and workload. It is not a universal hardware guarantee.

8.5 Confidential computing boundaries

An ordinary VM can keep another tenant out while leaving its own memory readable to the hypervisor. Protecting that memory from host administration requires enforcement below the software being excluded. Confidential computing protects data during computation using a hardware-backed, attested execution boundary. Memory encryption is common but is not required by every such design (Confidential Computing Consortium analysis20).

A trusted execution environment (TEE) is an execution area designed to protect the confidentiality and integrity of data and the integrity of executing code against specified outside parties. The selected boundary may enclose an application or an entire guest system. In the latter case, guest software remains trusted with the plaintext it processes (Confidential Computing Consortium analysis21).

For a concrete hardware design, AMD’s January 2020 white paper22 describes SEV-SNP, Secure Encrypted Virtualization with Secure Nested Paging. It uses VM-specific memory encryption with keys held in hardware beyond direct software reads. Checks of memory-page ownership and mappings resist a malicious hypervisor’s writes and remapping. Processor hardware and AMD Secure Processor firmware enforce these controls. The chip, that firmware, and the guest VM remain trusted. A host administrator therefore cannot recover private-memory plaintext through an ordinary host read. This is a design description, not a test of a current deployment.

Attestation supplies evidence about the selected system. It does not create its isolation. The Remote ATtestation procedureS (RATS) architecture23 separates the target environment from the environment that collects evidence about it and requires claims to stay securely associated with the target. Its evidence relationships do not establish the hardware protection described above.

Confidential-computing evidence has two levels. The Confidential Computing Consortium analysis24 supplies a general goal and threat-model limits and explicitly rejects absolute security. The exact CPU TEE, GPU mode, firmware, device assignment, attestation evidence, and key-release path must come from vendor and deployment documentation. Implementation assurance and current patch status remain specific to the deployed system.

Assessing a TEE requires its protected memory and execution boundary, excluded interfaces, trusted firmware, update process, measurement method, and relation to key release. An admission policy can require matching workload and firmware measurements and approved interfaces before making plaintext available. The result depends on what the selected platform actually measures and enforces. The host administrator sits outside the claimed area yet still controls scheduling, so denial of service stays possible. Model outputs and allowed logs leave through explicit interfaces that the boundary record must name.

A confidential VM alone does not extend that protection to an accelerator. The complete protected path is one versioned design. The design specifies the supported CPU TEE, the supported GPU model and confidential mode, guest and host software, firmware, and the passthrough arrangement. It also specifies I/O protection, CPU and GPU evidence, verifier policy, and key release. NVIDIA’s confidential-containers reference design25 describes CPU and GPU attestation with Kata utility VMs and Trustee-based key release, together with its own security limits. It is versioned vendor documentation rather than a general guarantee. Its living supported-platform matrix26 supplies the combinations claimed by a selected version, so the deployment record includes the matrix version and matching components.

Items crossing the boundary need their own record. Environment variables, Kubernetes ConfigMaps (objects that supply configuration values), mounted storage, network traffic, outputs, logs, and application code can remain untrusted, exposed, or able to disclose protected data. Composite attestation does not make those paths safe by itself.

Example

Tracing a protected memory read. Suppose W-4 runs inside the SEV-SNP guest described above, and the key service’s authenticated encrypted connection terminates inside that guest. Returned document text becomes plaintext in private guest memory. An ordinary host read cannot obtain that plaintext under the design’s assumptions. The application can still copy it into an output or log. SEV communication uses shared pages outside the private-memory guarantee, so encrypted transport must protect data crossing them (AMD design27). The boundary record identifies the image, private pages, interfaces, firmware, measurement, and decryption path. The host can still stop the workload.

Conclusion. Hardware mediates private-memory access, while guest code controls plaintext use and exit. The trace illustrates the claimed read boundary. It does not establish actual isolation, future state, side-channel resistance, availability, output confidentiality, or properties of unlisted firmware.

Timing, resource use, traffic shape, and shared hardware can leak information through side channels even when protected memory works as specified. Returning software to an earlier version can reintroduce defects the appraisal assumed fixed. Malicious or unmeasured firmware, an unmeasured device, or a defect in the trusted base can undercut the measurement the claim rests on. A platform statement therefore needs to place each of these inside or outside its threat model, name the mitigations it tested, and leave the rest as open review items. Protected execution also narrows hardware choice, reduces observability, and adds cost.

8.6 Remote attestation

A customer deciding at a distance whether to release data or keys can use cryptographic evidence about specified platform state instead of relying on broad assurance. In the decryption-session example, the key service requires acceptable evidence about the workload and platform before opening the session. Remote attestation is that evidence flow: an attester produces evidence about an environment, the verifier evaluates it against policy, and a relying party uses the result for an application decision (RATS architecture28). Here, freshness concerns the policy-dependent limit on how old attestation evidence may be. This differs from deciding which retrieved document version applies. State can change immediately after evidence is produced (RATS architecture29).

The verifier authenticates and appraises specified evidence against reference values and policy. The relying party then checks whether the resulting claims satisfy its release policy. A root of trust is a component accepted as a starting point for measurement or verification. A successful result supports only the listed claims under the accepted evidence, endorsements, reference values, freshness method, policy, and roots of trust. The architecture leaves evidence format, trust anchors, policy, and protected workload to each deployed system. A nonce (a one-use value), signature, or key-binding property therefore belongs to the selected evidence protocol rather than to attestation in general. The RATS architecture30 also states that it alone cannot show whether a deployed system mitigates its threats. The architecture allows one entity to combine roles, and it requires the design to assess the consequences of doing so. An appraisal result is further limited to the claims, policies, and accepted trust anchors the verifier actually used (RATS architecture31).

A reference value is a comparison value supplied for appraisal. An endorsement is an authenticated statement supporting trust in an attester’s capabilities, such as reliable evidence collection or signing. These have different jobs: one supplies an expected claim value. The other supports trusting how a claim was obtained or signed (RATS architecture32).

Suppose the selected deployment measures the launch image and configuration of W-4 by hashing their specified bytes. Its evidence includes those digests, the platform version, and the verifier’s challenge under a signature. The organization’s release team supplies reference set R-12 containing the digests of the approved image and configuration and the allowed platform versions. A manufacturer endorsement binds the evidence-signing public key to the protected signing capability. The verifier checks that statement against its accepted manufacturer trust anchor. These are assumed features of this example’s protocol, not mandatory fields in every attestation scheme. A changed configuration produces a different digest and fails R-12 under policy 5, even if the signature is valid. A matching digest identifies approved measured bytes. It does not show that those bytes are harmless or cover unmeasured state.

The verifier sends a fresh challenge to the attester, which returns authenticated evidence associating the challenge with claims about measured state. Reference values, endorsements, and appraisal policy support the verifier's attestation result. The relying party receives that result and applies application policy to allow, limit, or reject resource access.
Figure 8.2: This is one challenge-based attestation protocol for a key or credential release decision. The verifier checks evidence against reference values, endorsements, and policy. The relying party makes the separate access decision. Other attestation protocols can establish freshness differently. In the worked decryption-session example, the key service keeps raw key K-18 and returns authorized plaintext, so it does not perform the raw-key release pictured here. Attestation does not prove future state or harmless behavior.

Example

Following an attestation-gated decryption flow. Suppose raw document key K-18 remains inside the key service and approved workload W-4 requests a five-minute decryption session. In this example’s selected protocol, the verifier sends a fresh challenge and receives signed evidence associating that challenge with claims about the measured environment and workload. It checks the signature and freshness, compares the measurement with approved reference R-12, and applies policy version 5. The relying party checks the appraisal result and current document authorization, then opens the session at 10:03. During the session, W-4 sends an authorized decryption request and receives plaintext rather than raw key bytes. The retained records are the challenge, evidence, reference, policy, appraisal result, authorization decision, session decision, decrypt request, and output.

Conclusion. The session-opening decision was conditioned on one appraisal result and a current authorization check. The records do not establish future state, later plaintext use, or properties outside the measured boundary.

Trials of this flow seed approved, stale, and altered states, reporting accepted approved states alongside rejected stale or altered ones over all seeded cases. The machinery costs service dependencies, latency, reference-value maintenance, and a possible availability failure when the verifier cannot be reached. A compromised trust anchor, a rollback the policy accepts, or a state change after evidence collection can defeat the intended condition. Even a valid appraisal supports only the measured claim under the accepted roots of trust. The five-minute session is an assumed local design. Its deadline prevents new requests through that session. It does not erase plaintext, session credentials, logs, outputs, or raw key bytes exported by a different design.

Confidential computing and attestation do not fix application defects, authorize model actions, show that a model is harmless, validate output accuracy, protect every output or log, or guarantee availability. Approved measurements can identify approved bits or configuration claims. They do not establish that the approved model or code behaves safely. NVIDIA’s threat model for its confidential-computing VM design33 excludes malicious or vulnerable guest code, application logging, compromised attestation or key-release administration, side channels, physical attacks, and denial of service.

8.7 Provider audit scope

Provider paperwork gets the same scoping treatment. An audit scope states the systems, controls, locations, and time period examined. NCSC guidance34 asks providers to state which security aspects remain the user’s responsibility and where data may be used, accessed, or stored. A certificate or audit report therefore covers only its stated claim. A broad certificate, an expired report, or a configuration outside the examined scope creates false confidence if read as general proof. Deeper supplier evidence costs review time and may not be contractually available. The deployment owner therefore accepts, restricts, or rejects the platform after matching the evidence to the exact configuration and period. Unknowns stay recorded as unknowns rather than being assumed to be protected.

The platform record feeds the chapters that follow: release testing in release decisions (Chapter 14), live observation in system monitoring (Chapter 15), and incident recovery in recovery (Chapter 16). Placement choices use it in deployment choice (Chapter 19) and supplier duties continue in supplier responsibilities (Chapter 21). Reachable service interfaces are examined next in service exposure (Chapter 9).

Note

Chapter checkpoint. A rented-infrastructure model must read an encrypted document. The provider operates the hardware and firmware. The organization operates the workload and key policy. Which evidence flow can gate key release, and which claims stay open?

Answer. Record the workload and environment measurement, a fresh challenge, attester evidence, verifier trust anchors, approved reference values, appraisal policy and result, current document authorization, and the relying party’s session decision. In this design, the key service retains the raw key and opens a five-minute decryption session for the approved workload, checking authorization again for each request. Expiry blocks new decrypt requests through that session. It does not erase returned plaintext or copied session credentials. The flow can support the tested gate decision under the stated roots of trust. It does not establish the protected area’s isolation, later plaintext purpose, side-channel resistance, availability, output confidentiality, or future state.


  1. Shir Tamari and Sagi Tzadik, Wiz Research, “Wiz and Hugging Face Address Risks to AI Infrastructure,” Wiz blog, April 4, 2024, source. This was authorized research, and the report says the setup could have allowed cross-tenant access, not that any customer data was taken.↩︎

  2. Wiz Research, “Wiz Research Finds Critical NVIDIA AI Vulnerability Affecting Containers Using NVIDIA GPUs, Including Over 35% of Cloud Environments,” Wiz blog, September 26, 2024, source, with NVIDIA security advisory GHSA-q2v4-jw5g-9xxj, source. The later bypass is described in Yupeng (Roc), “CVE-2025-23359: Nvidia-container-toolkit: GPU Container Escape (CVE-2024-0132 fix bypass),” oss-security researcher disclosure, February 14, 2025, disclosure. These reports do not supply a prevalence estimate.↩︎

  3. Ramaswamy Chandramouli, Security Recommendations for Server-based Hypervisor Platforms, NIST SP 800-125A Rev. 1 (2018), abstract and Sections 1.1, 2, and 2.2.2 on baseline functions, threat sources, and device virtualization, source. Server-hypervisor guidance describing responsibilities and recommendations, not assurance for a named cloud or version.↩︎

  4. Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Sections 2.2 and 3, especially Section 3.5.2, source.↩︎

  5. Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Sections 4.3.4 and 4.4.2, printed pp. 23-25, source.↩︎

  6. Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Sections 3.3-3.5, printed pp. 15-18, source.↩︎

  7. Jerome H. Saltzer and Michael D. Schroeder (1975), “The Protection of Information in Computer Systems,” Proceedings of the IEEE 63(9), 1278-1308, section I.A.3, “Design Principles,” item f, “Least privilege,” source.↩︎

  8. Heidy Khlaaf and Tyler Sorensen, “LeftoverLocals: Listening to LLM Responses through Leaked GPU Local Memory,” Trail of Bits blog, January 16, 2024, exploit brief, testing across GPU platforms, and coordinated disclosure, source. Recovery from uninitialized GPU local memory on tested combinations with attacker access to a shared programmable GPU interface. Vendor and patch status are device and software specific.↩︎

  9. CERT/CC, VU#446598, “GPU kernel implementations susceptible to memory leak” (CVE-2023-4969), overview and vendor information, source. Tracks tested vendor status. It is not evidence that every GPU or memory region is affected.↩︎

  10. AMD, AMD-SB-6010, “GPU Memory Leaks,” CVE details and mitigation, source. The mode is disabled by default and requires an administrator to enable it on supported products. It is designed to prevent concurrent GPU processes and clear registers between processes. Performance can be affected.↩︎

  11. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), Introduction, “Who is responsible for developing secure AI?” PDF p. 7, source.↩︎

  12. Scott Rose, Oliver Borchert, Stu Mitchell, and Sean Connelly, Zero Trust Architecture, NIST SP 800-207 (2020), abstract and Section 2.1, “Tenets of Zero Trust,” source.↩︎

  13. Confidential Computing Consortium, A Technical Analysis of Confidential Computing, v1.3, updated November 2022, Section 5, “Threat Model,” source. General industry analysis that rejects absolute security. Assurance depends on the attacker and the deployment trust model.↩︎

  14. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), Secure deployment, “Make it easy for users to do the right things,” PDF p. 15, source.↩︎

  15. Elaine Barker, Recommendation for Key Management: Part 1, General, NIST SP 800-57 Part 1 Rev. 5 (2020), Sections 5-8, especially Section 8, source.↩︎

  16. Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan, Remote ATtestation procedureS (RATS) Architecture, IETF RFC 9334 (January 2023), Sections 3 and 4.1, source.↩︎

  17. Scott Rose, Oliver Borchert, Stu Mitchell, and Sean Connelly, Zero Trust Architecture, NIST SP 800-207 (2020), abstract and Section 2.1, source.↩︎

  18. Chris S. Lin, Joyce Qu, and Gururaj Saileshwar, “GPUHammer: Rowhammer Attacks on GPU Memories are Practical,” Proceedings of the Thirty-fourth USENIX Security Symposium (2025), abstract, threat model, evaluation, and Section 10, “Mitigations,” source. Demonstrated Rowhammer bit flips on an NVIDIA A6000 with GDDR6 under the paper’s co-location and attacker conditions; a bounded integrity example, not a claim about every GPU.↩︎

  19. Chris S. Lin, Joyce Qu, and Gururaj Saileshwar, “GPUHammer: Rowhammer Attacks on GPU Memories are Practical,” Proceedings of the Thirty-fourth USENIX Security Symposium (2025), abstract, threat model, evaluation, and Section 10, “Mitigations,” source. Demonstrated Rowhammer bit flips on an NVIDIA A6000 with GDDR6 under the paper’s co-location and attacker conditions; a bounded integrity example, not a claim about every GPU.↩︎

  20. Confidential Computing Consortium, A Technical Analysis of Confidential Computing, v1.3, updated November 2022, Sections 2.1-3.1, PDF pp. 5-6, source. General definition and hardware trust boundary, including designs that do not use memory encryption.↩︎

  21. Confidential Computing Consortium, A Technical Analysis of Confidential Computing, v1.3, updated November 2022, Sections 2.1-3.1, PDF pp. 5-6, source. General definition and hardware trust boundary, including designs that do not use memory encryption.↩︎

  22. Advanced Micro Devices (2020), AMD SEV-SNP: Strengthening VM Isolation with Integrity Protection and More, document 70366, January 2020, “Introduction,” “The Case for Integrity,” “Threat Model Details,” and “Reverse Map Table,” pp. 3-11, publication record, paper. A vendor design explanation whose cover explicitly disclaims a guarantee for resulting products. Current platform configuration, firmware, and mitigations require separate verification.↩︎

  23. Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan, Remote ATtestation procedureS (RATS) Architecture, IETF RFC 9334 (January 2023), Sections 3.1, 4.1, and 8.1, source. This is an Informational architecture, not a TEE isolation guarantee.↩︎

  24. Confidential Computing Consortium, A Technical Analysis of Confidential Computing, v1.3, updated November 2022, Section 5, “Threat Model,” source. General industry analysis that rejects absolute security. Assurance depends on the attacker and the deployment trust model.↩︎

  25. NVIDIA, Confidential Containers Reference Architecture, architecture overview and security considerations (living documentation), source. Versioned vendor design describing CPU and GPU attestation, Kata utility VMs, and Trustee-based key release; its own limits cover application defects, some control-plane inputs, unsafe storage mounts, most physical attacks, and availability.↩︎

  26. NVIDIA, Confidential Containers: Supported Platforms and Software Components, supported-platform matrix and GPU passthrough requirements (living documentation), source. Check the matrix for the software version selected for deployment.↩︎

  27. Advanced Micro Devices (2020), AMD SEV-SNP: Strengthening VM Isolation with Integrity Protection and More, document 70366, January 2020, “Introduction,” “The Case for Integrity,” “Threat Model Details,” and “Reverse Map Table,” pp. 3-11, publication record, paper. A vendor design explanation whose cover explicitly disclaims a guarantee for resulting products. Current platform configuration, firmware, and mitigations require separate verification.↩︎

  28. Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan, Remote ATtestation procedureS (RATS) Architecture, IETF RFC 9334 (January 2023), Sections 3 and 4.1, source.↩︎

  29. Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan, Remote ATtestation procedureS (RATS) Architecture, IETF RFC 9334 (January 2023), Section 10, “Freshness,” source.↩︎

  30. Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan, Remote ATtestation procedureS (RATS) Architecture, IETF RFC 9334 (January 2023), Section 12, especially Section 12.2, source.↩︎

  31. Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan, Remote ATtestation procedureS (RATS) Architecture, IETF RFC 9334 (January 2023), Sections 8-12, especially Sections 10 and 12.2, source. The RFC defines attestation roles and evidence flow. It does not certify a TEE or claim that attestation prevents side channels.↩︎

  32. H. Birkholz et al., Remote ATtestation procedureS (RATS) Architecture, RFC 9334 (2023), sections 4.1–4.2 and 8.1–8.3, source. The architecture distinguishes evidence, reference values, and endorsements. The digest fields and local approval policy in the example are design assumptions.↩︎

  33. NVIDIA, Confidential Computing reference architecture, “Trust and Threat Model,” responsibility table and what the design does and does not protect against, source. First-party design claim excluding malicious or vulnerable guest code, application logging, compromised attestation or key-release administration, side channels, physical attacks, and denial of service.↩︎

  34. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), Secure deployment, “Make it easy for users to do the right things,” PDF p. 15, source.↩︎