Part II: Attacks on data and models
Training data, requests, retrieved content, and stored copies expose different attack paths and require controls at different boundaries.
An attacker may alter training data, a request, a retrieved document, or the response returned to a user. Another attacker may try to obtain private data without changing it. Chapter 3 follows changes that reach model training or model files. Chapter 4 separates request and response tampering from adversarial inputs to a fixed model. Chapter 5 examines content admitted through search, and Chapter 6 examines transfers, stored copies, and information learned through queries. The affected object helps identify how harm may persist, where a control can act, and what evidence is useful. How long the harm lasts also depends on later use, retained copies, and recovery.
Chapters in this part
- Training poisoning and model tampering: Altering records used in training can change a model’s learned behavior. Investigating a suspected change requires evidence connecting the affected model version to its training data and training process.
- Input attacks and response tampering: Altering a request or its delivered response can produce a misleading result without changing model weights. Investigating that result requires records of what the caller submitted, what the model received and generated, and what the recipient obtained.
- RAG and search poisoning: Adding false information or hostile instructions to retrieved material can change an application’s answers without changing model weights. Stored copies can carry that influence into later requests.
- Data leakage and privacy attacks: An AI application can expose protected information through the data it sends, the copies it keeps, or the answers its model returns to specially chosen queries. Each route needs its own evidence, because a control on one route does not cover the others.