References
Cited sources retain their reference numbers. Broader background is listed separately as further reading.
The numbered entries identify sources cited in the guide. Further reading is listed separately. Footnotes connect a source to the particular claim and passage it supports. A source may offer guidance, a controlled experiment, an incident account or a legal text. Those roles are not interchangeable.
Sources cited
Apostol Vassilev, Alina Oprea, Alie Fordyce, Hyrum Anderson, Xander Davies, and Maia Hamin, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, National Institute of Standards and Technology (2025), title pages, abstract and sections 2.1 and 3.1, publication. The report supplies attack terminology. The book’s chapter order is a separate plan for the material. Additional cited coverage: Sections 2.3.1–2.3.2 cover availability, targeted, and clean-label poisoning (printed pp. 19–22. PDF pp. 32–35). Section 2.3.3 covers backdoors and limits of data filtering, trigger reconstruction, and model inspection (printed pp. 22–25. PDF pp. 35–38). Section 2.3.4 covers direct model changes and malicious collaborative updates (printed pp. 26–27. PDF pp. 39–40). Sections 3.2.1–3.2.3 discuss the generative-model supply chain (printed pp. 42–43. PDF pp. 55–56). These categories and scoped mitigation findings do not establish a universal poisoning detector or recovery guarantee. Report PDF. Sections 2.4.1 and 2.4.3, printed pp. 28-31 (PDF pp. 41-44), distinguish reconstruction, individual attributes and population properties. A representative reconstruction need not be an original record. Sections 2.1.1-2.1.4 and 3.1.1-3.1.3 are also cited for stage, goal, capability and knowledge. Section 3.1.1, printed pp. 36-38, distinguishes training and application integration; “Inference-time attacks,” item 4, printed p. 39, describes the agent cycle. InstructGPT is an example rather than a universal pipeline, and capabilities and permissions depend on the application. The overlapping application and service paths are this guide’s teaching map.
Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, National Institute of Standards and Technology (2023), sections 1 and 5, printed pp. 1–8 and 20–33, source. Relevant topics. AI system risk management and lifecycle context. The source also supports organizational risk processes. The framework organizes risk-management functions. Section 5, printed p. 21 (PDF p. 26), permits useful function order with governance in place and calls for iteration; section 5.1, printed pp. 21-24, describes GOVERN. Section 5.2, printed pp. 24-25, explains how MAP context informs measurement and management. Additional cited locators are Figure 4 and Section 3 for trustworthiness characteristics, Figure 5 and Sections 5.1-5.4 for Govern, Map, Measure and Manage, and Tables 2 and 4 for the cited context and risk decisions. The framework does not calculate an organization’s risk tolerance or approve a deployment.
National Institute of Standards and Technology (2024), Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, section 2.2, “Confabulation,” printed p. 6, and action MP-2.3-001, printed p. 24, source. This is a recommended evaluation action, not evidence that a particular checker is accurate. The profile lists generative-AI risks and suggested actions. Its introduction and section 3 are also cited. Action MP-1.1-001, printed p. 22, covers purpose mapping, fine-tuning and retrieval-based sources without ranking interventions universally.
UK National Cyber Security Centre, CISA, and international partners (2023), Guidelines for Secure AI System Development, version 1.0, November 27, section 1, “Design your system for security,” p. 10, and section 2, “Document your data, models and prompts,” p. 12, source. The guidance does not specify parser isolation, optical-character-recognition checks, metadata schemas, or version-selection rules. The executive summary, four guideline areas, and publishing and contributing organizations provide the front-matter context.
OWASP GenAI Security Project, OWASP Top 10 for LLM Applications 2026, version 2026, contents on PDF p. 3 and “LLM Top 10 at a Glance” on PDF p. 9, official publication and download. The table maps entries LLM01:2026 through LLM10:2026 to this guide; it is a navigation aid, not a claim of complete OWASP coverage. The retrieved PDF retains publication-date placeholders; the landing page is dated August 3 and the cover August 4, 2026.
Jerome H. Saltzer and Michael D. Schroeder, “The Protection of Information in Computer Systems,” Proceedings of the IEEE 63(9), 1278-1308 (1975), section I.A.3, Design Principles, especially complete mediation and least privilege, author-hosted text. Applying these principles to an AI application’s operations is the guide’s design reasoning.
Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy, “Explaining and Harnessing Adversarial Examples,” International Conference on Learning Representations (2015), author manuscript v3 (March 20, 2015). FGSM construction in Sections 2-4, especially Section 4 (PDF pp. 2-3), following the linear-model motivation in Section 3. Adversarial training in Sections 5-6 (PDF pp. 3-6). The expression is an untargeted first-order construction with parameters held fixed, not a proof of optimal attack or robustness. Version read.
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu, “Towards Deep Learning Models Resistant to Adversarial Attacks,” International Conference on Learning Representations (2018), author manuscript v4 (September 4, 2019). Robust optimization and PGD in Section 2 and Section 2.1 (PDF pp. 3-5). The inner optimization study follows in Section 3, and adversarial training and its experiments in Sections 4-5. The evaluated empirical resistance is not a universal certificate. Version read and conference record. Sections 3.1-3.3 connect repeated gradients and random restarts to the choice of PGD during training.
Anish Athalye, Nicholas Carlini, and David Wagner, “Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples,” International Conference on Machine Learning, PMLR 80:274-283 (2018). Section 3.1 identifies gradient-obfuscation symptoms, Section 4 develops adaptive attack techniques, and Section 5 with Table 1 reports the case study. Nine ICLR 2018 defenses were studied, with obfuscation identified in seven. Official publication.
Nicholas Carlini et al., “Extracting Training Data from Large Language Models,” Thirtieth USENIX Security Symposium, USENIX Association (2021), pp. 2633-2650, abstract and sections 4-6, official publication. The authors recovered and verified individual training examples from GPT-2. The result demonstrates possible disclosure under their procedure, not a universal extraction rate. Also cited are Sections 3-6 of the paper, especially candidate generation and ranking in Sections 4-5 and corpus verification in Section 6.1.
Patrick Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems 33 (2020), pp. 9459-9474, abstract, Figure 1, and Section 1, source. The paper combines a pretrained generator with a dense Wikipedia index and retriever. It does not prescribe the authorization design used here. Sections 2 and 4.5, “Index hot-swapping,” PDF pp. 7-8, discuss the retrieval/generator method and changing its Wikipedia index. These experiments do not establish the guide’s triage permissions or deployment choice.
Kai Greshake et al., “Not What You’ve Signed Up For: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” Proceedings of the Sixteenth ACM Workshop on Artificial Intelligence and Security (AISec) (2023), pp. 79–90, publication record. Author manuscript, arXiv:2302.12173v2, section 3, manuscript pp. 3–5, and section 4.2, pp. 6–10. These are demonstrations in the tested integrations and simulated applications, not evidence that every current integration has the same vulnerability.
Wei Zou, Runpeng Geng, Binghui Wang and Jinyuan Jia, “PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models,” Thirty-fourth USENIX Security Symposium, USENIX Association (2025), pp. 3827-3844, sections 3-5, official publication. Assumes insertion of malicious texts into the target database and evaluates specified retrievers and models.
Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song and Bo Li, “AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases,” Advances in Neural Information Processing Systems 37, 130185-130213 (2024), sections 3.2-3.3 and 4, conference paper. Core optimization assumes partial database writes and white-box embedder access. Transfer is evaluated separately. No additional model training is required.
Shawn Shan, Wenxin Ding, Josephine Passananti, Stanley Wu, Haitao Zheng, and Ben Y. Zhao. “Nightshade: Prompt-Specific Poisoning Attacks on Text-to-Image Generative Models.” 2024 IEEE Symposium on Security and Privacy. Version read: arXiv:2310.13828, first posted October 20, 2023. The authors present the method as a protective tool for artists, and the results concern the tested image models rather than any deployed service.
Nicholas Carlini, Matthew Jagielski, Christopher A. Choquette-Choo, Daniel Paleka, Will Pearce, Hyrum Anderson, Andreas Terzis, Kurt Thomas, and Florian Tramèr. “Poisoning Web-Scale Training Datasets is Practical.” 2024 IEEE Symposium on Security and Privacy (SP), pp. 407–425. Version read: arXiv:2302.10149v2, May 6, 2024. PDF locators refer to that author version: split-view poisoning, Sections 4.1–4.3, pp. 4–7 and Table 1; frontrunning mechanism, Section 3.3, p. 4, and Wikipedia timing experiments, Sections 5.1–5.4, pp. 8–11; perceptual-hash limits, Section 4.4, p. 7; trusted-reference and digest checks, Sections 6.1–6.2, p. 12. The collection experiments establish attack opportunities under those access and timing assumptions, not attack prevalence or universal downstream model effects.
Daniel Huynh and Jade Hardouin. “PoisonGPT: How We Hid a Lobotomized LLM on Hugging Face to Spread Fake News.” Mithril Security blog, July 9, 2023. Blog post. A vendor’s controlled demonstration, published alongside the vendor’s own product proposal.
Alexander Wan, Eric Wallace, Sheng Shen, and Dan Klein. “Poisoning Language Models During Instruction Tuning.” Proceedings of the Fortieth International Conference on Machine Learning, Proceedings of Machine Learning Research 202 (2023), pp. 35413–35425. Conference paper. Section 2, PDF p. 3, distinguishes clean-label and dirty-label access. Sections 4.1–4.2, PDF pp. 4–6, evaluate polarity attacks using the Tk-Instruct setup with T5 models. Sections 5–5.1, PDF pp. 6–7, test degenerate outputs on held-out tasks. Section 6.1, PDF pp. 7–8, and Figure 8, PDF p. 9, show filtering trade-offs and dependence on the checkpoint used to score examples. The filtering experiment uses a 3-billion-parameter model and 100 dirty-label examples. It does not establish the same effect for every filtering method or model family.
Jacob Steinhardt, Pang Wei Koh, and Percy Liang. “Certified Defenses for Data Poisoning Attacks.” Advances in Neural Information Processing Systems 30 (2017), pp. 3517–3529. Curran Associates. Conference paper. Abstract and Sections 2–3, PDF pp. 1–4, define the attacker budget, convex-loss setting, and approximate upper bounds. The analysis assumes concentration between training and test loss and little effect from removing clean-data outliers. It permits added poisoning examples under the specified constraints, not arbitrary changes to the learner or preprocessing code. This is not a general backdoor certificate for language models.
Evan Hubinger, Carson Denison, Jesse Mu, et al. “Sleeper Agents: Training Deceptive LLMs that Persist Through Safety Training.” arXiv:2401.05566, January 10, 2024. The backdoored models were built on purpose, and the study does not show that any deployed model is affected.
Peter Lee. “Learning from Tay’s introduction.” Official Microsoft Blog, March 25, 2016. Blog post. The company’s own short account, which does not describe the technical mechanism by which users changed the chatbot’s behavior.
Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. “Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!” International Conference on Learning Representations (ICLR), 2024. Conference paper. Section 4.1, PDF pp. 4–5, specifies Llama-2 Chat with 7 billion parameters, GPT-3.5 Turbo 0613, and a 330-prompt benchmark scored by a GPT-4 judge. Sections 4.2–4.3, PDF pp. 5–7, report adversarial fine tuning. Section 4.4, PDF p. 7, Table 3 and Figure 4, reports benign-data regressions and Llama-2 learning-rate/batch-size effects. Only training epochs were configurable through the GPT-3.5 fine-tuning API in the studied setup. These measured safety changes do not demonstrate an attack on the chapter’s constructed employee-feedback process.
Eric Rescorla, The Transport Layer Security (TLS) Protocol Version 1.3, RFC 8446 (2018). Design aims in Section 1 (pp. 6-7). Record protection in Section 5.2. Server authentication is part of the channel. Client authentication is optional. Standard: https://www.rfc-editor.org/rfc/rfc8446.
Steve Povolny and Shivangee Trivedi, “Model Hacking ADAS to Pave Safer Roads for Autonomous Vehicles,” McAfee Labs blog (February 19, 2020). Blog post. A vendor research demonstration on two cars with the first TACC implementation, with no road incident, and a 2020 car did not appear susceptible.
Fábio Perez and Ian Ribeiro (2022), “Ignore Previous Prompt: Attack Techniques For Language Models,” ML Safety Workshop, NeurIPS 2022, arXiv:2211.09527v1 (November 17, 2022), Sections 3-5 (PDF pp. 3-6), Table B11 (p. 14), and Table C4 (pp. 18-21), paper. The prompt-leaking experiments used
text-davinci-002and 35 public prompt templates. A result counted when the output contained the original instruction. Four runs per configuration gave condition-specific rates. This is controlled completion evidence, not validation of a commercial chatbot’s hidden prompt.AI Incident Database, “Incident 622: Chevrolet Dealer Chatbot Agrees to Sell Tahoe for $1,” incident dated December 18, 2023, based on an original post by Chris Bakke. Incident record. An incident summary built from public posts and news reports. The record describes a chatbot reply but does not document a completed sale.
OWASP GenAI Security Project, LLM01:2025 Prompt Injection, OWASP (2025), Description, Types, and Prevention and Mitigation Strategies, especially items 1, 3 and 6. Advisory guidance on instructions, filtering and separation, rather than measured effectiveness for a particular implementation. Official guidance. Also cited: Prevention and Mitigation Strategies, items 4-6. These are recommended practices, not controlled evaluations of effectiveness.
Alexander Wei, Nika Haghtalab, and Jacob Steinhardt, “Jailbroken: How Does LLM Safety Training Fail?,” Advances in Neural Information Processing Systems 36, 80079-80110 (2023). The two hypotheses appear in Sections 3.1-3.2. Sections 3-4 and Table 1 describe attacks, setup and results (PDF pp. 3-8). Appendix C.1 gives model details. Tests use 2023 GPT-4 and Claude v1.3 snapshots, and also evaluate GPT-3.5 Turbo, with bounded harmful-prompt sets. A published experiment, not an observed customer incident. Conference paper.
OWASP GenAI Security Project (2025), Top 10 for LLM Applications 2025, LLM05:2025, “Improper Output Handling,” description and “Common Examples of Vulnerability,” source. Listed outcomes are risk examples and do not prove a vulnerability in a particular application. Also cited: Prevention and Mitigation Strategies. These are recommendations, not measured effectiveness.
Liz Reid, “AI Overviews: About last week,” Google (May 30, 2024), source. Google’s own account attributes these examples to satirical and discussion-forum material. It also mentions searches apparently intended to elicit errors, so it does not establish that no user acted adversarially. It describes the feature at that time.
The Guardian, “ChatGPT search tool vulnerable to manipulation and deception, tests show” (December 24, 2024), as reported in TechCrunch, “ChatGPT Search can be tricked into misleading users, new research reveals” (December 26, 2024), source. The cited account is TechCrunch’s report of the Guardian tests, rather than a direct quotation from the Guardian article.
Johann Rehberger, “Spyware Injection Into Your ChatGPT’s Long-Term Memory (SpAIware),” Embrace The Red (September 20, 2024), source. A researcher demonstration. OpenAI’s fix in ChatGPT macOS app version 1.2024.247 closed the sending channel, not memory writes from untrusted content.
National Institute of Standards and Technology, NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, version 1.0 (2020), NIST CSWP 01162020, Appendix A, Table 2, ID.IM-P1-P8 and GV.PO-P1-P6, printed pp. 20 and 22, source. The inventory is separated from governing policies. These outcome categories do not form a complete AI data-classification scheme or legal permission test. Also cited: section 1 and Appendix A, Table 2, alternate official copy.
Hillai Ben-Sasson and Ronny Greenberg, “38TB of data accidentally exposed by Microsoft AI researchers,” Wiz blog (September 18, 2023), source. The researchers who found the exposure describe what the token made reachable, which is not a record of what others downloaded.
Bloomberg News, “Samsung Bans ChatGPT, Google Bard, Other Generative AI Use by Staff After Leak,” Bloomberg (May 2, 2023), source. A news report with no Samsung incident report behind it, and the nature of the material is reported rather than confirmed.
OpenAI, “March 20 ChatGPT outage: Here’s what happened,” OpenAI (March 24, 2023), source. The company’s own account of a bug in an application library, including its own estimate that actual exposure was extremely low.
Ramaswamy Chandramouli and Eric A. Hibbard (2025), Guidelines for Media Sanitization, NIST SP 800-88 Rev. 2, section 4.5, “Sanitization Assurance,” especially sections 4.5.1-4.5.2, printed pp. 24-25, source. Media sanitization does not by itself address logical copies in indexes, derived embeddings, supplier backups, or model weights.
Milad Nasr, Nicholas Carlini, Jonathan Hayase, et al., “Scalable Extraction of Training Data from (Production) Language Models,” arXiv:2311.17035 (November 28, 2023), paper, author explanation. Sections 4-5 of the original author PDF and the author explanation supply the cited evidence. The rate applies to the strongest tested configuration. For the closed model, the authors matched long output sequences against an auxiliary public-data collection, not a complete copy of the provider’s training corpus. The paper records disclosure to OpenAI on August 30, 2023.
Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song, “The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks,” Twenty-eighth USENIX Security Symposium (2019), pp. 267-284, Section 4, especially the exposure definition in Section 4.2, printed pp. 270-272, source. Exposure is defined relative to the canary format, space, model, and ranking rule, not every possible secret.
Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov, “Membership Inference Attacks Against Machine Learning Models,” IEEE Symposium on Security and Privacy (2017), pp. 3-18, DOI 10.1109/SP.2017.41, Sections IV-V, especially Section V-D and Figure 3 on known shadow membership and per-class attack training, source. The method uses full prediction vectors, true labels, and similarly trained shadow models. Its experiments do not establish a universal success rate or label-only equivalence. The cited author manuscript coverage is pp. 3-6; known shadow training and held-out splits supply the in/out labels, separate from task labels.
Matthew Fredrikson, Eric Lantz, Somesh Jha, Simon Lin, David Page, and Thomas Ristenpart, “Privacy in Pharmacogenetics: An End-to-End Case Study of Personalized Warfarin Dosing,” 23rd USENIX Security Symposium (2014), pp. 17-32, Sections 3.1-3.2 and Figure 2, printed pp. 20-23 (PDF pp. 5-8), conference paper. The attack uses known dose and background fields, black-box dose predictions, population frequencies, and model-error information. Its historical study does not establish general reconstruction success.
Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, et al., “Stealing Part of a Production Language Model,” arXiv:2403.06634 (March 11, 2024), source. Author version 1, Sections 3-5, 6, and 7.2 with Table 4. The production attack used top-five log-probabilities and controllable logit bias, recovering the final projection layer up to equivalent internal coordinates and the hidden size. The paper reports provider interface changes after disclosure. Sections 4.1-4.2 require rank and sufficiently varied hidden-vector conditions. Numerical precision and equivalent coordinates limit recovery. Table 4 reports historical layer costs of USD 4 and USD 12 for Ada and Babbage. This does not recover all weights or establish the same result for arbitrary text-only or current interfaces.
Florian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart, “Stealing Machine Learning Models via Prediction APIs,” Twenty-fifth USENIX Security Symposium (2016), pp. 601-618, Section 2 and Sections 4-6, source. The attacks study particular model classes and prediction services. Agreement alone does not establish identical parameters.
Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy (2014), Proposition 2.1, “Post-Processing,” printed p. 19, and Theorem 2.2, pure differential privacy for groups, printed p. 20, author manuscript. Post-processing concerns functions of the private mechanism’s output without a new access to protected data. Group privacy changes the protected relation and is distinct from composition of releases.
Martín Abadi et al., “Deep Learning with Differential Privacy,” ACM CCS (2016), pp. 308-318, DOI 10.1145/2976749.2978318, Algorithm 1 and Section 3.1, PDF p. 3, Section 3.2, PDF pp. 4-5, and Appendices A and B, PDF pp. 12-14, author manuscript v2, October 24, 2016. The reported privacy, compute, and accuracy results concern MNIST and CIFAR-10 experiments, with public CIFAR-100 pretraining for the CIFAR convolutional layers, not present-day large language model training. Sections 4-5 report the experiments. The add/remove record relation, fixed sampling rate and divisor, clipping bound, and sampling/accountant match govern the privacy calculation. Replacement and person-level relations require separate bounds.
Spracklen et al., “We Have a Package for You! A Comprehensive Analysis of Package Hallucinations by Code Generating LLMs,” Proceedings of the 34th USENIX Security Symposium (2025), source. The paper measures model output in a test setting and does not report packages installed or exploited in real organizations.
Murugiah Souppaya, Karen Scarfone, and Donna Dodson, Secure Software Development Framework (SSDF) Version 1.1: Recommendations for Mitigating the Risk of Software Vulnerabilities, NIST SP 800-218 (2022), Section 2 and task PS.3.2, source. Relevant topics: secure-development practices and software component records.
Hugging Face contributors, “Tokenizers,” Transformers v4.57.1 documentation, introduction and “Tokenizer classes,” source. The matching-vocabulary requirement concerns the model input representation, not assurance that a package is safe.
Hugging Face contributors, “Instantiate a big model,” Transformers v4.47.1 documentation, “Sharded checkpoints” and “Shard metadata,” source. The index maps parameter names to files. Shard sizes and file formats can vary.
Hugging Face, “Model Cards,” Hub documentation, “What are Model Cards?” and “Model card metadata,” retrieved September 24, 2026, source. Descriptive documentation is distinct from verification of the downloaded bytes and their behavior.
PyTorch Team, “Compromised PyTorch-nightly dependency chain between December 25th and December 30th, 2022,” PyTorch blog, December 31, 2022, source. The post does not say how many machines installed the package.
SLSA contributors, Supply chain Levels for Software Artifacts, specification v1.1, “Supply chain threats,” threat diagram and Build track, source. Version 1.1 is a fixed historical edition, marked retired on the current site.
Hugging Face, “Pickle Scanning,” Hub documentation (living page), “Hub’s Security Scanner” and “Disclaimer,” source. Retrieved September 16, 2026. The documentation describes scanning and warns that it is not foolproof. It does not measure attack frequency.
JFrog Security Research, “Data Scientists Targeted by Malicious Hugging Face ML Models with Silent Backdoor,” JFrog blog, February 27, 2024, source. JFrog judges most payloads to be researchers’ work, so the count is not a count of attacks.
ReversingLabs, “Malicious ML models discovered on Hugging Face platform,” RL Blog, February 6, 2025, and the related press release, source. ReversingLabs calls the two models likely proof of concept and reports no victims.
PyTorch contributors, “torch.load,” PyTorch 2.8 documentation, Warning below the parameter descriptions, source.
PyTorch contributors, “Serialization semantics,” PyTorch 2.8 documentation, updated May 19, 2025, “torch.load with weights_only=True,” source. This description restricts object construction and on-demand imports. It is not an assurance of complete loader isolation.
Wiz Research, “Probllama: Ollama Remote Code Execution Vulnerability (CVE-2024-37032) – Overview and Mitigations,” Wiz blog, June 24, 2024, source. The exposed-instance count measures internet reachability, not vulnerable versions or confirmed exploitation.
Seth Larson, “Supply-chain attack analysis: Ultralytics,” The Python Package Index Blog, December 11, 2024, source. The analysis comes from the package index and covers the publishing path, not the effect on users who installed the versions.
Amazon Web Services, “Malicious script injected into Amazon Q Developer for Visual Studio Code (VS Code) Extension,” security advisory GHSA-7g7f-ff96-5gcw, July 2025, source. This is the vendor’s own account, and its statement that the code could not run and affected no customers is disputed in press coverage.
SLSA contributors, Supply chain Levels for Software Artifacts, specification v1.1, “Security levels,” “Build L1” and “Build L2,” source. Build L1 requires a build record. Build L2 adds signed records from a hosted platform. This is not a model-behavior guarantee. The build record’s established name is described in “Provenance,” Purpose and Model, same fixed edition.
Shir Tamari and Sagi Tzadik, Wiz Research, “Wiz and Hugging Face Address Risks to AI Infrastructure,” Wiz blog, April 4, 2024, source. This was authorized research, and the report says the setup could have allowed cross-tenant access, not that any customer data was taken.
Wiz Research, “Wiz Research Finds Critical NVIDIA AI Vulnerability Affecting Containers Using NVIDIA GPUs, Including Over 35% of Cloud Environments,” Wiz blog, September 26, 2024, source. Researcher report, not a prevalence estimate.
NVIDIA, security advisory GHSA-q2v4-jw5g-9xxj, source. Retrieved September 19, 2026. Vendor security record.
Yupeng (Roc), “CVE-2025-23359: Nvidia-container-toolkit: GPU Container Escape (CVE-2024-0132 fix bypass),” oss-security researcher disclosure, February 14, 2025, disclosure. The disclosure does not supply a prevalence estimate.
Ramaswamy Chandramouli, Security Recommendations for Server-based Hypervisor Platforms, NIST SP 800-125A Rev. 1 (2018), abstract and Sections 1.1, 2, and 2.2.2 on baseline functions, threat sources, and device virtualization, source. Server-hypervisor guidance describing responsibilities and recommendations, not assurance for a cloud service or software version.
Murugiah Souppaya, John Morello, and Karen Scarfone, Application Container Security Guide, NIST SP 800-190 (2017), Sections 2-4, source. Relevant topics: images, registries, orchestration, containers, host systems, and countermeasures.
Heidy Khlaaf and Tyler Sorensen, “LeftoverLocals: Listening to LLM Responses through Leaked GPU Local Memory,” Trail of Bits blog, January 16, 2024, exploit brief, testing across GPU platforms, and coordinated disclosure, source. Recovery from uninitialized GPU local memory on tested combinations with attacker access to a shared programmable GPU interface. Vendor and patch status are device and software specific.
CERT/CC, VU#446598, “GPU kernel implementations susceptible to memory leak” (CVE-2023-4969), overview and vendor information, source. Retrieved September 16, 2026. Tracks tested vendor status. It is not evidence that every GPU or memory region is affected.
AMD, AMD-SB-6010, “GPU Memory Leaks,” CVE details and mitigation, source. Retrieved September 16, 2026. The mode is disabled by default and requires an administrator to enable it on supported products. It is designed to prevent concurrent GPU processes and clear registers between processes. Performance can be affected.
Scott Rose, Oliver Borchert, Stu Mitchell, and Sean Connelly (2020), Zero Trust Architecture, NIST SP 800-207, section 2.1, tenet 3, printed p. 6, source. The publication gives an architecture principle. It does not define retrieval filters or show that filtering after retrieval prevents disclosure.
Confidential Computing Consortium, A Technical Analysis of Confidential Computing, v1.3, updated November 2022, Section 5, “Threat Model,” source. General industry analysis that rejects absolute security. Assurance depends on the attacker and the deployment trust model. Sections 2.1-3.1, PDF pp. 5-6, define the hardware trust boundary and include designs without memory encryption.
Elaine Barker, Recommendation for Key Management: Part 1, General, NIST SP 800-57 Part 1 Rev. 5 (2020), Sections 5-8, especially Section 8, source.
Henk Birkholz, Dave Thaler, Michael Richardson, Ned Smith, and Wei Pan, Remote ATtestation procedureS (RATS) Architecture, IETF RFC 9334 (January 2023), Sections 8-12, especially Sections 10 and 12.2, source. The RFC defines attestation roles and evidence flow. It does not certify a TEE or claim that attestation prevents side channels.
Chris S. Lin, Joyce Qu, and Gururaj Saileshwar, “GPUHammer: Rowhammer Attacks on GPU Memories are Practical,” Proceedings of the Thirty-fourth USENIX Security Symposium (2025), abstract, threat model, evaluation, and Section 10, “Mitigations,” source. Demonstrated Rowhammer bit flips on an NVIDIA A6000 with GDDR6 under the paper’s co-location and attacker conditions; a bounded integrity example, not a claim about every GPU.
NVIDIA, Security Notice on Rowhammer, reading pointer. The notice body and publication date have not been verified in the source set checked here. The chapter’s ECC mitigation discussion instead follows Lin, Qu, and Saileshwar (2025), Section 10.
Advanced Micro Devices (2020), AMD SEV-SNP: Strengthening VM Isolation with Integrity Protection and More, document 70366, January 2020, “Introduction,” “The Case for Integrity,” “Threat Model Details,” and “Reverse Map Table,” pp. 3-11, publication record, paper. A vendor design explanation whose cover explicitly disclaims a guarantee for resulting products. Current platform configuration, firmware, and mitigations require separate verification.
NVIDIA, Confidential Containers Reference Architecture, architecture overview and security considerations (living documentation), source. Retrieved September 16, 2026. Versioned vendor design describing CPU and GPU attestation, Kata utility VMs, and Trustee-based key release; its own limits cover application defects, some control-plane inputs, unsafe storage mounts, most physical attacks, and availability.
NVIDIA, Confidential Containers: Supported Platforms and Software Components, supported-platform matrix and GPU passthrough requirements (living documentation), source. Retrieved September 16, 2026. Supported combinations depend on the documentation and deployment versions.
NVIDIA, Confidential Computing reference architecture, “Trust and Threat Model,” responsibility table and what the design does and does not protect against, source. Retrieved September 16, 2026. First-party design claim excluding malicious or vulnerable guest code, application logging, compromised attestation or key-release administration, side channels, physical attacks, and denial of service.
Oligo Security, “ShadowRay: First Known Attack Campaign Targeting AI Workloads Exploited In The Wild,” Oligo blog, March 26, 2024, source. The almost $1 billion figure is Oligo’s estimate of the value of machines and compute that might have been compromised, and Ray’s developers dispute that CVE-2023-48022 is a vulnerability.
National Institute of Standards and Technology, “CVE-2025-3248 Detail,” National Vulnerability Database, published April 7, 2025, source. Retrieved September 19, 2026. The entry records the flaw and its catalog status, not the extent of real attacks.
Trend Micro Research, “Critical Langflow Vulnerability (CVE-2025-3248) Actively Exploited to Deliver Flodrix Botnet,” Trend Micro, June 17, 2025, source. This is one security vendor’s observation of one campaign, not a count of affected Langflow servers.
Wiz Research, “Wiz Research Uncovers Exposed DeepSeek Database Leaking Sensitive Information, Including Chat History,” Wiz blog, January 29, 2025, source. The report describes open databases found by researchers, not code execution or confirmed misuse by attackers.
Kubernetes contributors, “Security Checklist,” Kubernetes documentation (living page), “Network security” and “Authentication and authorization,” source. Retrieved September 16, 2026. This living checklist must be applied to the deployed Kubernetes version.
Hugging Face, “Space secrets leak disclosure,” Hugging Face blog, May 31, 2024, source. The disclosure does not identify the attacker, the method, or the number of secrets accessed.
Kubernetes contributors, “Service Accounts,” Kubernetes documentation (living page), “What are service accounts?” and “Grant permissions to a ServiceAccount,” source. Retrieved September 16, 2026.
Ramaswamy Chandramouli and Zack Butcher, A Zero Trust Architecture Model for Access Control in Cloud-Native Applications in Multi-Location Environments, NIST SP 800-207A (2023), Sections 2-3, especially ID-SEG-REC-2 and ID-SEG-REC-3, source. Relevant topics: service identity and service-to-service authorization recommendations.
Oligo Security, “ShadowRay 2.0: Active Global Campaign Hijacks Ray AI Infrastructure Into Self-Propagating Botnet,” Oligo blog, November 18, 2025, source. The server count measures exposure, Oligo says AI-generated code is only strongly implied, and the September 2024 start date is only possible.
Aim Labs (Itay Ravia), “Breaking down ‘EchoLeak’, the First Zero-Click AI Vulnerability Enabling Data Exfiltration from Microsoft 365 Copilot,” Cato Networks (originally Aim Security), current republication reviewed September 28, 2026, source. The researchers’ own account, republished after Cato acquired Aim. The original Aim Labs page is no longer available. It gives the attack chain but no disclosure timeline, CVE number, or CVSS score.
Bill Toulas, “Zero-click AI data leak flaw uncovered in Microsoft 365 Copilot,” BleepingComputer, June 11, 2025, source. A press report and the only source here for the January 2025 report date and May 2025 fix.
Microsoft Security Response Center, “CVE-2025-32711: M365 Copilot Information Disclosure Vulnerability,” Microsoft, June 11, 2025, source. The vendor record reports a CVSS 3.1 base score of 9.3 and
Exploited: No, with a brief vulnerability description and a mitigation FAQ. The NVD record separately lists NIST’s CVSS 3.1 base score as 7.5. These fields were checked September 29, 2026 in MSRC’s CVRF data and NVD.Sizhe Chen, Julien Piet, Chawin Sitawarin, and David Wagner (2025), “StruQ: Defending Against Prompt Injection with Structured Queries,” Thirty-fourth USENIX Security Symposium, 2383-2400, sections 3.1, 4.2-4.4, 5.1, and 6, published paper. Sections 4.2-4.4, PDF pp. 8-10, connect filtering, delimiter token identifiers, and clean/attacked instruction tuning. Table 2 on PDF p. 11 includes successful optimization attacks. The attacker controls the data segment. The guide’s delivery-day trace is illustrative.
JFrog Security Research (2024), “When Prompts Go Rogue: Analyzing a Prompt Injection Code Execution in Vanna.AI,” JFrog Blog, source. Researcher account of one library design. The researcher account records a maintainer response with a hardening guide. Exposure depends on deployment settings.
National Vulnerability Database, “CVE-2024-5565,” published May 31, 2024, record. Retrieved September 19, 2026. Vulnerability record accompanying the researcher report. It does not establish the settings of a particular deployment.
Johann Rehberger (2023), “OpenAI Begins Tackling ChatGPT Data Leak Vulnerability,” Embrace The Red, December 20, 2023, source. The reporter’s own tests of a partial, web-first change. It does not measure how complete the mitigation was or how it changed later.
Edoardo Debenedetti et al., “Defeating Prompt Injections by Design,” arXiv:2503.18813v2, June 24, 2025, sections 2-5 and 7-9, especially 5.1-5.4 on models, policies, value metadata, and dependency tracking, inspected manuscript. The author publication record lists IEEE SaTML 2026. The inspected text is the dated manuscript. Trusted requests and uncompromised memory are assumptions. Policy mistakes and side channels remain limits. The payroll example is a proposed application policy.
PromptArmor (2024), “Data Exfiltration from Slack AI via Indirect Prompt Injection,” PromptArmor, August 20, 2024, source. A researcher demonstration against one product at one point in time. It shows a working path, not its use by attackers.
Austin Larsen, Matt Lin, Tyler McLellan, and Omar ElAhdan (2025), “Widespread Data Theft Targets Salesforce Instances via Salesloft Drift,” Google Cloud Blog, Google Threat Intelligence Group, August 26, 2025, source. An investigator’s account of stolen OAuth credentials. No model or agent behavior was involved.
Model Context Protocol contributors (2026), Model Context Protocol Specification, revision 2026-07-28, “Overview” and “Security and Trust & Safety,” source. Relevant topics. Protocol roles, resources, prompts, tools, message schemas, authorization, consent, and security considerations. “Key Details” and “Base Protocol” distinguish host, client and server roles. Discovery is separate from business authorization.
Model Context Protocol contributors (2026), Model Context Protocol Specification, revision 2026-07-28, “Authorization,” “Purpose and Scope,” “Protocol Requirements,” “Roles,” and “Token Requirements,” source. The profile is optional and concerns HTTP deployments that implement it. STDIO uses environment credentials. Business identity and object permissions remain application responsibilities. The Introduction is also cited.
Michael B. Jones and Dick Hardt. The OAuth 2.0 Authorization Framework: Bearer Token Usage. RFC 6750, Internet Engineering Task Force (IETF), October 2012. Official specification. Bearer use requires no proof of a separate cryptographic key. Applies to OAuth bearer deployments.
T. Lodderstedt (editor), S. Dronia, and M. Scurtescu. OAuth 2.0 Token Revocation. RFC 7009, Internet Engineering Task Force (IETF), August 2013. Official specification. Revocation requests can propagate with delay, and revoking only the refresh token does not promise immediate invalidation of issued access tokens.
Justin Richer (editor), OAuth 2.0 Token Introspection, RFC 7662, IETF (October 2015), sections 2.2 and 4, official specification. The authorization server supplies token state and metadata. A cached response may not reflect a later revocation.
Torsten Lodderstedt, John Bradley, Andrey Labunets, and Daniel Fett. Best Current Practice for OAuth 2.0 Security. RFC 9700 / BCP 240, Internet Engineering Task Force (IETF), January 2025. Official specification. Minimal privileges with audience, action, and resource restrictions (Section 2.3). Refresh-token rotation or sender constraint for public clients (Sections 2.2.2 and 4.14.2).
Open Policy Agent contributors. Open Policy Agent (OPA), documentation overview (undated), introduction and “Writing Policies with Rego.” Official documentation. Retrieved 2026-09-16. OPA evaluates structured input against declarative Rego policies and returns decisions. The surrounding application must enforce them at the operation boundary. The chapter’s payment fields and approval rule are constructed examples, not built-in OPA behavior or evidence that a payment integration has been tested.
Model Context Protocol contributors (2026), Model Context Protocol Specification, revision 2026-07-28, “Tools,” “Data Types,” “Tool,” annotations requirement, source.
Marco Milanta and Luca Beurer-Kellner (2025), “GitHub MCP Exploited: Accessing private repositories via MCP,” Invariant Labs Blog, May 26, 2025, source. A demonstration with no CVE. It shows a possible path, not an observed attack.
Postmark (2025), “Information Regarding Malicious ‘postmark-mcp’ Package,” Postmark Blog, September 25, 2025, source. Statement by the service imitated by the package.
Ravie Lakshmanan (2025), “First Malicious MCP Server Found Stealing Emails in Rogue Postmark-MCP Package,” The Hacker News, September 29, 2025, reporting Koi Security’s findings, source. Koi Security’s findings are cited through a news report rather than Koi’s own publication.
OWASP GenAI Security Project (2025, December 9), “OWASP Top 10 for Agentic Applications: The Benchmark for Agentic Security in the Age of Autonomous AI,” release announcement, “The OWASP Agentic Top 10,” source. Relevant topics. Agent goal hijacking, tool misuse, identity and privilege misuse, memory poisoning, and related agent risks.
Kevin Townsend (2026), “Capsule Security Launches ‘AI Circuit Breaker’ to Stop Rogue Agents,” SecurityWeek, September 3, 2026, source. A news report of vendor claims. The figures come from Capsule’s own evaluation and are not independently replicated.
D. Fett, B. Campbell, J. Bradley, T. Lodderstedt, M. Jones, and D. Waite. OAuth 2.0 Demonstrating Proof of Possession (DPoP). RFC 9449, Internet Engineering Task Force (IETF), September 2023. Official specification. Access-token protection requires an issued DPoP token with matching proof validation. A Bearer response does not acquire that protection.
Michael B. Jones, Anthony Nadalin, Brian Campbell (editor), John Bradley, and Chuck Mortimore. OAuth 2.0 Token Exchange. RFC 8693, Internet Engineering Task Force (IETF), January 2020, Sections 1.1, 2.1-2.3, 4.1, and Appendix A.2, official specification. Subject and actor inputs, issuer policy, target token, independent lifetimes, and current-actor claims are distinct parts of the exchange. Section 1.1 distinguishes delegation and impersonation. Section 4.1 limits decisions to top-level claims and the current actor, not nested actor history. Input and output token lifetimes and revocation are not automatically linked.
Michael B. Jones, John Bradley, and Nat Sakimura (2015), JSON Web Token (JWT), RFC 7519, IETF, May 2015, Sections 2-3 and 7.2, source. A representation and validation specification, not a grant of application permission.
Model Context Protocol contributors (2026), Authorization Security Considerations, revision 2026-07-28, “Token Audience Binding and Validation” and “Access Token Privilege Restriction,” source. Token passthrough is forbidden. A separate token issued by the downstream API’s authorization server is required. This does not require every deployment to implement RFC 8693. “Token Theft” and “Confused Deputy Problem” are also cited in this revision. The repository text is an alternate presentation of these sections.
Beatrice Nolan (2025), “AI-powered coding tool wiped out a software company’s database in ‘catastrophic failure’,” Fortune, July 23, 2025, source. News coverage of a first-person account, with no forensic report behind it.
Hammond Pearce et al., “Asleep at the Keyboard? Assessing the Security of GitHub Copilot’s Code Contributions,” 2022 IEEE Symposium on Security and Privacy, pp. 754–768, publication record. Author manuscript, arXiv:2108.09293v3, abstract and sections IV–VI, especially “Threats to Validity,” PDF pp. 1 and 4–13. The experiments used the Copilot version, prompts, languages, and weakness scenarios described in the paper. They do not establish the security of current coding assistants.
Artem Chaikin and Shivan Kaul Sahib (2025), “Agentic Browser Security: Indirect Prompt Injection in Perplexity Comet,” Brave Blog, August 20, 2025, source. Written by a company that sells a competing browser, about one attack path and its own retest.
Adam Barth (2011), The Web Origin Concept, RFC 6454, IETF, sections 8.1-8.3, pp. 14-16, source.
Google, “Image understanding,” Gemini API documentation, “Passing images to Gemini” and “Passing inline image data,” retrieved September 24, 2026, source. Documents an image-plus-text API input, not the service’s complete internal representation or resistance to prompt injection. The evidence-record comparison is an engineering consequence of which intermediate values the application actually receives.
Hugging Face (2026), “Anatomy of a Frontier Lab Agent Intrusion: A Technical Timeline of the July 2026 Incident,” Hugging Face Blog, July 27, 2026, source. The affected company’s own reconstruction of the incident, not an independent investigation.
Jiaqi Luo et al. (2026), “Autonomy Comes with Costs: Detecting Denial-of-Service Vulnerabilities Caused by Resource Abusing in LLM-based Agents,” Proceedings of the USENIX Security Symposium 2026, 3991–4010, sections 2.1–2.4, 4.1.1, 6.1 and 9, official publication and paper, Limit: studied selected open-source agent web applications with source-code-guided fuzzing and did not estimate prevalence for arbitrary external model services or compromised runtimes.
Edoardo Debenedetti, Jie Zhang, Mislav Balunovic, Luca Beurer-Kellner, Marc Fischer and Florian Tramèr (2024), “AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents,” Proceedings of the Thirty-eighth Conference on Neural Information Processing Systems, Datasets and Benchmarks Track, sections 3, 3.1, 3.4 and Figure 1, proceedings paper, Limit: evaluation environment with specified tasks and attacker goals, not a deployment incident rate. See Chapter 14 for task completion, attacker goals and evaluation limits.
UK AI Security Institute, Frontier AI Trends Report (web report, n.d.), sections on cyber capabilities, chemistry and biology, autonomy skills, loss of control, safeguards, and evaluation cheating, official report. Retrieved September 16, 2026. Bounded evaluations, not a general release guarantee.
OpenAI, Preparedness Framework, version 2 (2025), sections 2.2, “Tracked Categories,” and 2.3, “Research Categories,” printed pp. 4-8, versioned framework. A provider policy, not measured effectiveness or a universal release standard.
Drew Keller et al. (2026), Expanding the AI Evaluation Toolbox with Statistical Models, NIST AI 800-3, February 2026, section 3.1.1, “Choosing an Accuracy Estimand,” printed p. 9, and sections 3.2-3.3.1, printed pp. 9-11, source.
Cynthia Dwork et al. (2015), “Generalization in Adaptive Data Analysis and Holdout Reuse,” Advances in Neural Information Processing Systems 28, abstract and Section 1, proceedings PDF pp. 1-3, paper. Only the design lesson is used here, which is to search adaptively, then freeze and estimate on untouched cases. It is not the paper’s formal reusable-holdout guarantees, which belong to its stated algorithms and sampling assumptions.
NIST/SEMATECH (online edition, retrieved 2026-09-16), e-Handbook of Statistical Methods, section 7.2.4.1, “Confidence Intervals,” Wilson and exact methods, source. The binomial formulas assume a fixed number of independent trials with a common probability. One-sided and two-sided confidence levels differ.
Lawrence D. Brown, T. Tony Cai, and Anirban DasGupta, “Interval Estimation for a Binomial Proportion,” Statistical Science 16, no. 2 (2001), Introduction, pp. 101-102. Section 3.1.1, equation (4), pp. 107-108. Section 4.2.1, pp. 113-114. Section 5, “Concluding Remarks,” p. 115, paper. Wilson recommended for small samples over the poor-coverage Wald interval. Clopper-Pearson guarantees coverage at or above the nominal level but is conservative. Binomial model conditions apply: fixed n, one common success probability. The published paper is also available from coauthor T. Tony Cai. DOI 10.1214/ss/1009213286.
Karen Kent and Murugiah Souppaya (2006), Guide to Computer Security Log Management, NIST SP 800-92, sections 2.1-2.3, 4.1-4.4, and 5.1-5.2, especially 5.1.2, “Log Storage and Disposal,” source.
Alex Nelson, Sanjay Rekhi, Murugiah Souppaya, and Karen Scarfone, Incident Response Recommendations and Considerations for Cybersecurity Risk Management: A CSF 2.0 Community Profile, NIST SP 800-61 Rev. 3 (2025), Executive Summary and section 3, Tables 2–3, source. Relevant topics. Incident response integrated with cybersecurity risk management. Table 3, RC.RP-03 through RC.RP-06, especially RC.RP-05 R1-R2, printed p. 34 (PDF p. 42), calls for remedying root causes before production restoration and verifying restored operation. The guide’s URL-fetch removal and retests are an invented implementation example.
Coalition for Content Provenance and Authenticity (2025), C2PA Technical Specification, version 2.2, May 2025, sections 1.2-1.3, 14.2-14.3, and 15.12, source. C2PA supplies signed statements and validation states. The organization then determines whether those statements support the proposed use. Sections 2.3.13, 2.4.3 and 9.3.1 describe signed claims and content binding, distinct from statistical watermark detection.
Kelley Dempsey et al. (2011), Information Security Continuous Monitoring (ISCM) for Federal Information Systems and Organizations, NIST SP 800-137, section 3.1, subsection “Security Impact Analysis,” and sections 3.5-3.6, printed pp. 33-35, source.
OpenAI (2025), “Sycophancy in GPT-4o: What happened and what we’re doing about it,” OpenAI, April 29, source. Provider account of its own release.
OpenAI (2025), “Expanding on what we missed with sycophancy,” OpenAI, May 2, source. The company’s own follow-up account, with no independent review of the evaluations it describes.
Karen Kent, Suzanne Chevalier, Tim Grance, and Hung Dang (2006), Guide to Integrating Forensic Techniques into Incident Response, NIST SP 800-86, sections 3 and 3.1, especially 3.1.2, “Collecting the Data,” source.
UK National Cyber Security Centre (2025, May 7), Impact of AI on Cyber Threat from Now to 2027, NCSC assessment, “Assessment,” and “AI impact on stages of cyber intrusion to 2027,” source.
South China Morning Post (2024, May 17), “UK multinational Arup confirmed as victim of HK$200 million deepfake scam that used digital version of CFO to dupe Hong Kong employee,” South China Morning Post, source. The report relies on police and company statements and gives no technical account of how the synthetic media was made.
AI Incident Database, “Incident 983: Scammers Reportedly Used AI Voice Clone and YouTube Footage to Impersonate WPP CEO in Unsuccessful Scam Attempt,” AI Incident Database, incident reported May 10, 2024, source. The entry collects news reports and adds no independent investigation of the call.
Federal Bureau of Investigation (2025, May 15), Senior U.S. Officials Impersonated in Malicious Messaging Campaign, Alert I-051525-PSA, “Recommendations,” “Spotting a Fake Message,” and “How to Protect Yourself from Potential Fraud or Loss of Sensitive Information,” source.
Wiz (2025), “s1ngularity: supply chain attack leaks secrets on GitHub: everything you need to know,” Wiz Blog, August 27, 2025 (updated August 29), source. The post describes the malware’s behavior and has no controlled comparison with a search that did not use AI tools.
Anthropic (2025, November 13), “Disrupting the first reported AI-orchestrated cyber espionage campaign,” Anthropic, source. This is the model provider’s report. Its observations and interpretation are attributed to Anthropic rather than treated here as an independently reconstructed account.
MITRE (live knowledge base, retrieved 2026-09-16), Adversarial Threat Landscape for Artificial-Intelligence Systems (ATLAS), technique and case-study records. Official data schema, source. Relevant topics. AI-related adversary tactics, techniques, mitigations, case studies, and evidence maturity. MITRE ATLAS contributors, ATLAS Data, repository README, “Distributed ATLAS Data” and “Data Format.” Retrieved September 28, 2026, from repository README. This README describes the data types and format. It does not refresh the September 16 technique/case-record citation or verify current catalogue contents. Catalogue presence alone establishes neither local vulnerability nor prevalence.
Anthropic (2026, April 7), “Claude Mythos Preview’s cybersecurity capabilities,” Anthropic, source. The model’s developer ran and reported the tests. The source set cited here contains no independent replication of this comparison. The Firefox comparison and footnote 1 concern JavaScript shell exploits in a Firefox 147 content-process test harness without the browser process sandbox or other defense layers. The 181 successful developments are not a total-attempt denominator, a count of distinct new flaws, or an operational browser compromise rate. The engine vulnerabilities were patched in Firefox 148.
OpenAI (2024, May 30), “Disrupting deceptive uses of AI by covert influence operations,” OpenAI, source. The ratings are OpenAI’s own assessment of activity it chose to disclose. The paragraph immediately before “Attacker trends” supplies the 1-6 Breakout Scale and OpenAI’s category-2 interpretation for the five disclosed operations. It does not measure belief change independently.
John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein (2023), “A Watermark for Large Language Models,” Proceedings of ICML, PMLR 202, 17061-17084, Sections 2-3.1, Algorithm 2 and equation (3), PDF pp. 2-4; Section 5, “Private Watermarking,” PDF p. 5; Section 6, experiments; and Appendix B, “Evaluating Repetitive Text,” PDF pp. 18-19, source. The experiments concern the proposed detector and OPT-family models, including OPT-6.7 B. Detection and quality results depend on the tested generation settings and transformations. Participating generation applies the token-subset bias; detection needs matching tokenizer and partition settings and, in private mode, a key or detector API. Short, predictable, repeated or edited text limits inference.
Rich Piazza, Emily Ratliff, Stephan Relitz, and Christian Studer, eds., STIX Version 2.1 Errata 01, OASIS Standard incorporating Draft 01 of Errata 01 (April 2, 2025), section 3.2, “Common Properties,” confidence property, and Appendix A, “Confidence Scales,” source.
Google Project Zero, “From Naptime to Big Sleep: Using Large Language Models To Catch Vulnerabilities In Real-World Code,” Project Zero blog (November 1, 2024), source. The post reports one finding by the tool’s own developers and does not measure how often the tool finds real flaws.
DARPA, “AI Cyber Challenge marks pivotal inflection point for cyber defense,” DARPA news (August 8, 2025), source. The counts come from the organizer’s own competition scoring, mostly on synthetic vulnerabilities, and do not show patch quality in production code.
OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026: LLM01 PromptInjection, final repository source, revision 7e144e05142b2105faa1603178a1163a2a34df24, description and prevention or mitigation sections, source. Risk descriptions and recommendations, not measurements of prevalence or control effectiveness.
XBOW, “The Road to Top 1: How XBOW Did It,” XBOW blog (June 24, 2025), source. The company reports its own results, and the report counts show submission volume more than validated findings.
Industrial Cyber News Desk, “Dream’s Hero scores 96.6% on CyberGym to rank first globally in autonomous cybersecurity research benchmark,” Industrial Cyber (September 9, 2026), source. The article reports a score from an evaluation run on the vendor’s own infrastructure, with no independent check.
CrowdStrike, “CrowdStrike Launches the Charlotte AI AgentWorks Ecosystem for Building Secure Agents,” press release (March 25, 2026), source. A vendor press release describes intended features and does not state the contract terms, operators, or update duties for each model.
Open Source Initiative, The Open Source AI Definition, version 1.0 (2024), “What is Open Source AI” and “Preferred form to make modifications to machine-learning systems,” source, The definition is a community standard. A legal conclusion about a selected license requires separate analysis.
FinOps Foundation, Unit Economics, living framework (n.d.), “Definition,” “Define Unit Metrics which support Organizational Goals,” and maturity level “Run,” source, retrieved September 16, 2026. Actual costs depend on workload, supplier, and date.
FinOps Foundation, FinOps for AI, living framework (n.d.), “FinOps Considerations for AI,” source, retrieved September 16, 2026.
Cherilyn Pascoe, Stephen Quinn, and Karen Scarfone, The NIST Cybersecurity Framework (CSF) 2.0, NIST CSWP 29 (2024), section 1, printed pp. 1–2, and Appendix A, pp. 15–23, source. Relevant topics. Cybersecurity risk governance and outcome categories across an organization.
European Parliament and Council, Regulation (EU) 2024/1689 of June 13, 2024 (Artificial Intelligence Act), OJ L, 2024/1689 (July 12, 2024), consolidated text dated July 27, 2026, consolidation. Cited provisions are Articles 2, 3, 6, 11-14, 16, 25-26, 86, 111 and 113, and Annex III; exact paragraph locators and role, classification, territorial and transition conditions are in the chapter notes. The consolidation is a documentation tool. The published Official Journal acts control. Regulation (EU) 2026/1744 supplies the cited amendments. The Chapter III dates do not delay Chapter IX Article 86, whose own thresholds, exceptions and transition conditions apply. The Commission Article 86 reproduction supplies original-text context. The Article 3 reproduction supplies the original systemic-risk definition. Neither reproduction establishes the amended edition.
European Parliament and Council, Regulation (EU) 2026/1744 of July 8, 2026 (Digital Omnibus on AI), OJ L, 2026/1744 (July 24, 2026), Article 1(39)-(40) and Article 4, published act. Article 4 provides entry into force on the third day after publication, July 27, 2026.
British Columbia Civil Resolution Tribunal, Moffatt v. Air Canada, 2024 BCCRT 149 (February 14, 2024), source. The CRT explains that its decisions do not set binding precedent for CRT members in later cases. This dispute does not establish a general legal rule for AI answers under other laws. Historical case facts here are reported by Lisa R. Lifshitz and Roland Hung, “BC Tribunal Confirms Companies Remain Liable for Information Provided by AI Chatbot,” Business Law Today (February 2024), published commentary, “The Tribunal’s Decision in Moffatt.” This is secondary case reporting. The tribunal decision body was not available for this review.
Garante per la protezione dei dati personali, “ChatGPT, il Garante privacy chiude l’istruttoria. OpenAI dovrà realizzare una campagna informativa di sei mesi e pagare una sanzione di 15 milioni di euro,” press release (December 2024), source. Retrieved September 24, 2026. The notice records that Rome court judgment No. 4153/2026, published March 18, 2026, upheld the appeal and that the authority temporarily removed decision No. 755 from its website. This confirms a later court development, but the judgment’s full reasoning was not inspected.
ANSA, “Tribunale Roma annulla multa da 15 milioni di euro del Garante Privacy a OpenAI,” ANSA (March 20, 2026), source. This is a news report of the ruling, not the ruling itself, and the court’s full reasoning was not reviewed.
U.S. Department of Health and Human Services, Summary of the HIPAA Security Rule, reviewed August 7, 2026, headings Introduction, General Rules and Administrative Safeguards, source, retrieved September 24, 2026. Relevant topics. Safeguards and risk management for electronic protected health information. The August 7, 2026 regulator review distinguishes the Security Rule in effect from proposed changes. Coverage concerns covered entities, business associates and electronic protected health information, not all health data or clinical safety.
ISO and IEC, ISO/IEC 42001:2023, Information technology - Artificial intelligence - Management system, first edition (2023), public ISO overview, “What is ISO/IEC 42001?,” source, This citation covers the public scope description, not the licensed normative clauses. Relevant topics. AI management system requirements and continual organizational improvement.
Katherine Schroeder, Hung Trinh, and Victoria Yan Pillitteri, Measurement Guide for Information Security: Volume 1, Identifying and Selecting Measures, NIST SP 800-55v1 (2024), section 3.1.1, printed pp. 12–14, and sections 3.3.1–3.3.4, pp. 17–18, source.
OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026: LLM02 SensitiveInformationDisclosure, final repository source, revision 7e144e05142b2105faa1603178a1163a2a34df24, description and prevention or mitigation sections, source. Risk descriptions and recommendations, not measurements of prevalence or control effectiveness.
OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026: LLM03 ExcessiveAgency, final repository source, revision 7e144e05142b2105faa1603178a1163a2a34df24, description and prevention or mitigation sections, source. Risk descriptions and recommendations, not measurements of prevalence or control effectiveness.
OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026: LLM07 Misinformation, final repository source, revision 7e144e05142b2105faa1603178a1163a2a34df24, description and prevention or mitigation sections, source. Risk descriptions and recommendations, not measurements of prevalence or control effectiveness.
OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026: LLM09 VectorAndEmbeddingWeaknesses, final repository source, revision 7e144e05142b2105faa1603178a1163a2a34df24, description and prevention or mitigation sections, source. Risk descriptions and recommendations, not measurements of prevalence or control effectiveness.
European Parliament and Council, Regulation (EU) 2016/679 of April 27, 2016 (General Data Protection Regulation), OJ L 119 (May 4, 2016), pp. 1-88, Article 4(4), published act. The profiling definition is cited through AI Act Article 3(52). This use does not establish the classification or compliance of the guide’s hypothetical payment deployment.
OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026: LLM06 UnboundedConsumption, final repository source, revision 7e144e05142b2105faa1603178a1163a2a34df24, description and prevention or mitigation sections, source. Risk descriptions and recommendations, not measurements of prevalence or control effectiveness.
Michael Fagan et al., IoT Device Cybersecurity Capability Core Baseline, NISTIR 8259A (2020), section 2, Table 1: “Data Protection,” printed p. 7. “Logical Access to Interfaces,” p. 8. “Software Update,” p. 9. “Cybersecurity State Awareness,” p. 10, source.
H. Brendan McMahan et al., “Communication-Efficient Learning of Deep Networks from Decentralized Data,” Proceedings of the 20th International Conference on Artificial Intelligence and Statistics, PMLR 54 (2017), pp. 1273-1282, source, Objective equation (1) and client weighting on PDF p. 3. Section 2 on PDF p. 4 and Algorithm 1 on PDF p. 5 for local epochs, batch size, client fraction, and weighted averaging. Printed Algorithm 1 sums over \(K\) after selecting \(S_t\) and does not define excluded clients. The guide’s sampled-client example uses the explicit choice \(a_{k,t}=n_k/\sum_{j\in S_t}n_j\).
Keith Bonawitz et al., “Practical Secure Aggregation for Privacy-Preserving Machine Learning,” Proceedings of the 2017 ACM SIGSAC Conference on Computer and Communications Security, pp. 1175-1191, publication record, conference paper. Sections 3 to 5 and Figure 4 specify the cryptographic assumptions, masking, secret sharing, and selective recovery. Section 6.1, Theorem 6.3, printed p. 1182, gives honest-but-curious input privacy for \(c<t\). Sections 6.2 to 6.3, Theorem 6.5, printed p. 1184, give active input privacy for \(2t>n+c\), with at least \(t-c\) honest inputs in the disclosed sum. Active security additionally needs authenticated signing keys, signatures, a consistency check, random oracle and Two Oracle Diffie-Hellman assumptions. Section 5 gives completion under honest dropouts when at least \(t\) clients survive. Section 7 evaluates the honest-but-curious variant. The active result does not guarantee correctness or availability. The full-proof author manuscript, IACR ePrint 2017/281 uses different page and theorem numbering. In that author manuscript, masking/setup appears in Sections 4.0.2-4.0.3 and 5, passive Theorem 6.3 on p. 8, and active Appendix A Theorem A.2 on p. 18. Section 7 reports passive performance and dropout simulations. These measurements are not an evaluation of the active protocol.
Peva Blanchard et al., “Machine Learning with Adversaries: Byzantine Tolerant Gradient Descent,” Advances in Neural Information Processing Systems 30 (2017), Definition 1 and score \(s(i)\) on p. 4 and Proposition 1 on p. 5, section 5 Proposition 2 on p. 6, source, The resilience result needs \(2f+2<n\), independent unbiased estimates, bounded variance, and \(\eta(n,f)\sqrt{d}\sigma<\|g\|\), plus separate convergence conditions. The selected vector may be Byzantine.
Dong Yin et al., “Byzantine-Robust Distributed Learning: Towards Optimal Statistical Rates,” Proceedings of the 35th International Conference on Machine Learning, PMLR 80 (2018), pp. 5650-5659, Definitions 1 to 2 on PDF p. 3 and Algorithm 1 and Assumptions 1 to 4 on PDF p. 4 and Theorems 1 to 6 on PDF pp. 5 to 6, source, Coordinate median and trimmed mean need distinct moment, smoothness, and fraction assumptions and a common distribution empirical gradient model, not multi epoch non identical updates.
Eugene Bagdasaryan et al., “How To Backdoor Federated Learning,” Proceedings of the 23rd International Conference on Artificial Intelligence and Statistics, PMLR 108 (2020), pp. 2938-2948, abstract, section 3, section 4.2, and section 5, source, The evaluated tasks use CIFAR-10 images and Reddit word prediction with attacker control of participating clients and their submitted updates.
Keith Stouffer et al., Guide to Operational Technology (OT) Security, NIST SP 800-82 Rev. 3 (2023), section 2.3.1, printed pp. 11-12. section 4.1.2, Table 5, p. 52. section 5.3.1, p. 79, source.
Stav Cohen, Ron Bitton, and Ben Nassi, “Here Comes The AI Worm: Unleashing Zero-click Worms that Target GenAI-Powered Applications,” arXiv:2403.02817v1 (March 5, 2024), versioned paper. Section 1, “Ethical Considerations,” limits the experiments to authors’ laboratory applications. Sections 3.2 to 3.3 distinguish model replication from application propagation. Section 4.1, Figure 1 and steps 1 to 10, traces stored correspondence, later retrieval, replication into a generated reply, sending, and contamination of the recipient’s retrieval store. Section 5 describes the separate application-flow variant. The proposed gate and matched tests are illustrative designs, not measured defenses from this paper.
UK Department for Science, Innovation and Technology, Frontier AI Safety Commitments, AI Seoul Summit 2024, updated February 7, 2025, Outcome 1, commitments I-V, and footnotes 1 and 3, source.
European Commission, General-Purpose AI Code of Practice (2025), “The 3 chapters of the code,” Safety and Security paragraph, source, retrieved September 16, 2026.
PyTorch contributors, PyTorch documentation 2.14,
torch.nn.Linear, “Variables,” versioned documentation. This layer lists learnable weight and bias separately.Hugging Face contributors, Transformers documentation 5.17.0, “Parameter-efficient fine-tuning,” introduction and “Training,” versioned documentation. This example updates adapter parameters while keeping the base model frozen.
Long Ouyang et al., “Training language models to follow instructions with human feedback,” Advances in Neural Information Processing Systems 35 (2022), Figure 2 and section 3.1, conference paper. The procedure concerns InstructGPT and the study’s data and evaluations.
Hugging Face contributors, Transformers documentation 5.17.0, “Chat templates,” introduction and “Using apply_chat_template,” versioned documentation. The described formatting is implementation-specific.
Hugging Face contributors, Transformers documentation 5.17.0, “Tool use,” “Passing tools” and “Tool-calling Example,” versioned documentation. Caller software handles the generated request and appends the result.
Microsoft, Threat Modeling Tool threats, STRIDE model table. Retrieved September 28, 2026, from article. The accessible table supports the six categories, not exhaustive AI threat discovery.
Bruce Schneier, “Attack Trees,” Dr. Dobb’s Journal (December 1999), “Enter Attack Trees” and “Creating Attack Trees,” author-hosted article. This citation supports the described goal decomposition and AND/OR branches. It does not credit this article with every predecessor or provide measurements for this guide’s system.
U.S. Department of Health and Human Services, 45 CFR 160.103, definitions of individually identifiable health information, protected health information and electronic protected health information, including PHI paragraphs (1)-(2) and electronic media (1)-(2), definition text, and 45 CFR 164.302, Security Rule applicability, applicability text, eCFR display through September 24, 2026. The display is authoritative but unofficial. The education-record exclusion concerns FERPA-covered records. The student-treatment exclusion refers to 20 USC 1232g(a)(4)(B)(iv). The employment-record exclusion requires a covered entity acting as employer. Record category and the regulated entity’s role require separate checks. This recorded display does not establish later legal changes, entity-specific classification or complete compliance.
Shengye Wan et al., “CyberSecEval 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models,” arXiv:2408.01605v2 (September 6, 2024), section 3.2 and Appendix A.1, paper.
AICPA (2022), AU-C Section 9402, “Audit Considerations Relating to an Entity Using a Service Organization,” interpretation No. 1, paragraph .04, PDF pp. 2-3, interpretation. This financial-statement audit interpretation describes SOC 2 reports and customer controls. It does not establish any particular supplier’s scope or performance.
OpenAI (2026, July 21), “OpenAI and Hugging Face partner to address security incident during model evaluation,” “What happened during this incident” and the July 28 update, operator account. Read alongside Hugging Face’s reconstruction in reference 122. These are company accounts of the evaluation incident, not independent forensic validation.
Further reading
The survey covers federated optimization, privacy, security, and open research problems. It is retained as further reading, not presented as support for a cited claim in the current chapters.
- Peter Kairouz et al., “Advances and Open Problems in Federated Learning,” Foundations and Trends in Machine Learning 14(1-2) (2021), pp. 1-210, publication record. Author manuscript, arXiv:1912.04977, abstract and introductory overview. Relevant topics. Federated learning methods, privacy, security, and open problems.