Orientation and reading paths
Shared system settings and reading paths connect the security questions across the book.
Abstract
AI security connects attacks on data and models to the software, people, and organizations that use them. This study guide examines those connections across internal and external models, document retrieval, and applications that can act on a user’s behalf. Its main sources combine public security guidance with research on specific attacks and defenses.
NIST supplies the book’s main vocabulary for attacks and risk management. Apostol Vassilev et al. from government, universities, and industry wrote its Adversarial Machine Learning1 taxonomy. It distinguishes attack goals, attacker access and knowledge, and the stage at which an attack affects a system. Elham Tabassi’s AI Risk Management Framework2 connects risk assessment to organizational responsibility, while NIST’s Generative AI Profile3 applies that framework to generative systems. These reports provide terminology and management guidance, not proof that a particular deployment is secure.
The UK National Cyber Security Centre (NCSC), the US Cybersecurity and Infrastructure Security Agency (CISA), and international partner agencies published the Guidelines for Secure AI System Development4, with contributions from companies and research organizations. Their guidance covers secure design, development, deployment, operation and maintenance. It connects AI-specific concerns to the software, suppliers and operational processes around a model.
The OWASP GenAI Security Project5 contributes community guidance on application risks, including prompt injection, information disclosure and excessive authority. Its Top 10 documents help identify failure paths and possible controls. A risk category or suggested mitigation does not establish an attack’s frequency or a control’s effectiveness in an organization’s own application.
Some security principles predate modern AI. In their 1975 paper, Jerome Saltzer and Michael Schroeder6 describe limiting a program’s privileges and checking access to protected objects. Those principles remain relevant when a model helps an application select an operation: generated text does not itself grant permission to carry it out.
Research papers supply methods and evidence for more specific questions. Ian Goodfellow et al.7 developed a practical method for constructing misleading model inputs. Aleksander Madry et al.8 studied training against permitted adversarial changes, while Anish Athalye et al.9 showed how unsuitable tests could make selected defenses appear stronger than they were. Carlini et al.10 later extracted memorized training text from GPT-2 under specified query and verification procedures. These studies establish mechanisms and bounded results, not failure rates for every deployed model.
Research on retrieval connects model behavior to external data. Patrick Lewis et al.11 combined a language generator with a searchable document collection. Kai Greshake et al.12 demonstrated that instructions in external content could redirect tested applications. Wei Zou et al.13 examined malicious texts inserted into a retrieval database in their PoisonedRAG study, and Zhaorun Chen et al.14 examined poisoned agent memory or knowledge stores in their AgentPoison study. Their attacker access, models and tasks differ. In particular, influencing a request through retrieved content does not by itself show that the model’s trained weights have changed.
The chapter order, recurring examples, and diagrams form this guide’s own plan for the material. A framework recommendation and an experimental result answer different questions, so the explanations distinguish recommended practice from demonstrated behavior and retain the conditions that limit each conclusion.
Audience and system context
Suppose an employee sends a confidential incident report to an AI support assistant and requests a summary. The application retrieves an internal document and sends selected text to an external model. It returns the answer to the employee. A similar application can use a model operated inside the organization. It can also suggest a ticket change for a person to review, or submit a change request that the receiving service should check against the employee’s permissions.
The visible answer alone does not reveal the data path or the employee’s document rights. The model instructions and application authority also remain unknown. AI security analysis connects the system setting to attacker access, possible harm, controls, and evidence.
Highly technical readers who are relatively new to AI are the primary audience. The guide assumes familiarity with software systems, identity, networks, data handling, and ordinary security controls. It explains the AI concepts needed for each security question when they first matter.
Shared system settings
The guide reuses the settings below to keep the security questions concrete. Each setting specifies the organization, users, data, model operator, application capability, and attacker access before the control analysis.
- Employee support assistant. Employee identity, private documents, an application, and an internal or external model remain stable. Data paths, retrieval, model instructions, action authority, and incident state change.
- Controlled model laboratory. A fixed training or evaluation task provides controlled access to data, model parameters, model files, and compute. The attack mechanism, attacker access, defense, and test method change.
- Security analyst assistant. A read-only security task uses synthetic evidence, a human decision, and a fixed service need. Attacker or defender use, measurement method, model choice, and operating cost change.
- Organization portfolio. The same system map, threat record, control evidence, cost record, and responsibility record apply across multiple systems. Supplier relationships, sector duties, and technical assumptions change.
The deployment boundary and application capability are separate choices. A model may run inside the organization or on a managed platform. An external provider may operate it instead. The surrounding application may answer or retrieve documents. It may also propose an action or carry out an authorized action. These choices create different data paths and duties.
Book map
Data checks, platform protections and authorization checks address different steps of an AI request. Their results can be combined only when they refer to the same system and intended use. Part 1 establishes that shared description, including the attacker’s access and the outcomes the organization intends to prevent.
The flow then branches into three questions: how attacks through data can change model behavior or expose information in Part 2, how the platform protects the running service in Part 3, and how untrusted text can influence an application or its actions in Part 4. These branches meet in Part 5, where tests and operating records show how the controls work together and what happens when they fail.
Parts 6 and 7 use that evidence to assess AI in cybersecurity and to make organizational decisions. Part 8 asks which conclusions need to be revisited when the system or its setting changes. Changed assumptions lead back to the system description and to the tests that depended on them.
Book contents
The chapters below follow the full reading order. The learning paths that follow offer shorter routes for different roles.
Part 1: System and threat modeling
Part 2: Attacks on data and models
- Chapter 3: Training poisoning and model tampering
- Chapter 4: Input attacks and response tampering
- Chapter 5: RAG and search poisoning
- Chapter 6: Data leakage and privacy attacks
Part 3: Platform and service compromise
- Chapter 7: Supply chain compromise
- Chapter 8: Isolation on shared infrastructure
- Chapter 9: Runtime intrusion and resource abuse
Part 4: Untrusted instructions and excess authority
- Chapter 10: Indirect prompt injection
- Chapter 11: Retrieval access and disclosure
- Chapter 12: Agent action authorization
- Chapter 13: Limits on agent execution
Part 5: Security evaluation and response
- Chapter 14: Evidence for release decisions
- Chapter 15: Live monitoring and operating limits
- Chapter 16: AI incidents: containment and recovery
Part 6: AI in attack and defense
- Chapter 17: AI-assisted attacks
- Chapter 18: Evaluating AI for defense
- Chapter 19: Comparing AI deployment options
Part 7: Organizational AI governance
- Chapter 20: Shadow AI and unmanaged use
- Chapter 21: Responsibility and legal duties
- Chapter 22: Prioritizing security investment
Part 8: Security under changing conditions
- Chapter 23: Adapting controls across sectors
- Chapter 24: Reassessing security after technology changes
- Chapter 25: Capstone: defending a system decision
Appendices and reference material
Learning paths
The routes share Part 1, then emphasize the questions most relevant to a reader’s work. System builders need the data, platform and application branches. Cybersecurity practitioners focus on AI-assisted attacks and defensive use. Security leaders need the technical evidence that supports organizational decisions.
The routes are:
- AI system security: Parts 1 through 5, followed by Part 8.
- AI in cybersecurity: Parts 1, 4, 5, 6, and 8. Before Part 4, read the retrieval workflow and source checks, direct prompt injection and jailbreaks, and retention and deletion evidence in Part 2.
- Security leadership and governance: Parts 1, 2, 5, 7, and 8.
- Complete study route: Parts 1 through 8 in order.
The shorter routes support the listed roles and leave some deployment work to other readers. The whole-system capstone needs the system and threat records, control evidence, a deployment and cost comparison, and a responsibility record. Readers on a shorter route can complete the relevant missing teaching in Parts 2, 3, 4, 6, or 7, or obtain those records from teammates with that preparation. Each applicable attack branch still needs a control and test. A branch judged irrelevant needs a reason tied to the case facts.
Scope of the guide
The guide concerns the security of AI systems and the security effects of using AI. It introduces model training, generation, retrieval, and agents only to explain the relevant attack and control mechanisms. Its scope excludes foundation model development and general cloud administration, as well as the full lifecycle for developing secure software.
Legal and regulatory duties differ by jurisdiction and sector. The chapters identify where those duties change a security decision without offering legal advice, and examples state their assumptions without claiming that one control fits every organization or deployment.
Evidence and references
Research findings, measured results, historical attributions, formal definitions, and standards based processes are cited beside the claims they support. Sources favor established standards bodies, security organizations, university material, reputable technical publications, and influential peer reviewed research.
Common industry practices may appear without a citation when the text identifies them as common practice rather than a measured fact. The text states open questions, disputed results, and limits of available evidence directly. Footnotes support the claim near them. They do not replace the explanation in the main text.
The glossary appendix collects the defined terms in one alphabetical reference.
Apostol Vassilev, Alina Oprea, Alie Fordyce, Hyrum Anderson, Xander Davies, and Maia Hamin, Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, National Institute of Standards and Technology (2025), title pages, abstract and sections 2.1 and 3.1, publication. The report supplies attack terminology. The book’s chapter order is a separate plan for the material.↩︎
Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, National Institute of Standards and Technology (2023), sections 1 and 5, publication. NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, National Institute of Standards and Technology (2024), introduction and section 3, publication. The framework organizes risk-management functions. The profile lists generative-AI risks and suggested actions.↩︎
Elham Tabassi, Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1, National Institute of Standards and Technology (2023), sections 1 and 5, publication. NIST, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1, National Institute of Standards and Technology (2024), introduction and section 3, publication. The framework organizes risk-management functions. The profile lists generative-AI risks and suggested actions.↩︎
NCSC, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), executive summary and four guideline areas, guidelines and publishing and contributing organizations. Lifecycle guidance rather than experimental evidence of a control’s effectiveness.↩︎
OWASP GenAI Security Project, OWASP GenAI LLM Top 10 2026, OWASP (August 3, 2026), introduction and risk entries, project publication. Community guidance on application risks and mitigations, not a measured frequency distribution of attacks. Chapters that discuss an earlier edition identify its year explicitly.↩︎
Jerome H. Saltzer and Michael D. Schroeder, “The Protection of Information in Computer Systems,” Proceedings of the IEEE 63(9), 1278-1308 (1975), section I.A.3, Design Principles, especially complete mediation and least privilege, author-hosted text. Applying these principles to an AI application’s operations is the guide’s design reasoning.↩︎
Ian J. Goodfellow, Jonathon Shlens and Christian Szegedy, “Explaining and Harnessing Adversarial Examples,” International Conference on Learning Representations (2015), author manuscript v3, section 4, PDF pp. 2-3, version read. Introduces the fast gradient sign method. It does not establish an optimal attack against every model.↩︎
Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras and Adrian Vladu, “Towards Deep Learning Models Resistant to Adversarial Attacks,” International Conference on Learning Representations (2018), author manuscript v4 (September 4, 2019), sections 2-5, version read. Studies optimization for attack resistance and adversarial training under specified perturbation sets. Empirical resistance is not a universal guarantee.↩︎
Anish Athalye, Nicholas Carlini and David Wagner, “Obfuscated Gradients Give a False Sense of Security: Circumventing Defenses to Adversarial Examples,” International Conference on Machine Learning, PMLR 80:274-283 (2018), sections 3-5 and Table 1, official publication. Examines selected ICLR 2018 defenses under their specified threat models.↩︎
Nicholas Carlini et al., “Extracting Training Data from Large Language Models,” Thirtieth USENIX Security Symposium, USENIX Association (2021), pp. 2633-2650, abstract and sections 4-6, official publication. The study extracts and verifies memorized text from GPT-2. It does not estimate leakage across all language models.↩︎
Patrick Lewis et al., “Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks,” Advances in Neural Information Processing Systems 33, 9459-9474 (2020), sections 2 and 4.5, conference paper. Describes a generator coupled to an external document index, including an index-update experiment. It does not establish retrieval as a security control.↩︎
Kai Greshake, Sahar Abdelnabi, Shailesh Mishra, Christoph Endres, Thorsten Holz and Mario Fritz, “Not what you’ve signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection,” arXiv:2302.12173v2 (May 5, 2023), sections 3-4, version read. Demonstrations on synthetic applications and selected 2023 systems support the mechanism, not a claim about every current application.↩︎
Wei Zou, Runpeng Geng, Binghui Wang and Jinyuan Jia, “PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models,” Thirty-fourth USENIX Security Symposium, USENIX Association (2025), pp. 3827-3844, sections 3-5, official publication. Assumes insertion of malicious texts into the target database and evaluates specified retrievers and models.↩︎
Zhaorun Chen, Zhen Xiang, Chaowei Xiao, Dawn Song and Bo Li, “AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases,” Advances in Neural Information Processing Systems 37, 130185-130213 (2024), sections 3.2-3.3 and 4, conference paper. Core optimization assumes partial database writes and white-box embedder access. Transfer is evaluated separately. No additional model training is required.↩︎