Securing AI Systems

Author

Vitaly Rubinovich

Compilation date
Publisher ID
study-guide-publisher@0.8.14+873a57a

A security engineer checks a brass and glass AI apparatus with controlled document intake, a separated action arm, gauges, and a paper trace recorder.

This study guide explains how to protect AI systems across data, models, applications, operations, and organizational decisions. Its main sources are the US National Institute of Standards and Technology (NIST), the UK National Cyber Security Centre (NCSC), and international partner agencies. It also draws on the OWASP application security community and research papers examining specific attacks, defenses, and their limits.

Contents

Front matter

Part I: System and threat modeling

A system map and a threat model give technical and management readers the same description of the system, its operators, attacker access, and possible harm.

  • Chapter 1: Mapping the AI system
    A system map records the people, components, data, permissions, and evidence that a security decision about an AI assistant depends on.
  • Chapter 2: Threat modeling and controls
    A threat record connects an attacker’s access to a possible failure, the controls expected to prevent it, and the evidence needed to test them.

Part II: Attacks on data and models

Training data, requests, retrieved content, and stored copies expose different attack paths and require controls at different boundaries.

  • Chapter 3: Training poisoning and model tampering
    Altering records used in training can change a model’s learned behavior. Investigating a suspected change requires evidence connecting the affected model version to its training data and training process.
  • Chapter 4: Input attacks and response tampering
    Altering a request or its delivered response can produce a misleading result without changing model weights. Investigating that result requires records of what the caller submitted, what the model received and generated, and what the recipient obtained.
  • Chapter 5: RAG and search poisoning
    Adding false information or hostile instructions to retrieved material can change an application’s answers without changing model weights. Stored copies can carry that influence into later requests.
  • Chapter 6: Data leakage and privacy attacks
    An AI application can expose protected information through the data it sends, the copies it keeps, or the answers its model returns to specially chosen queries. Each route needs its own evidence, because a control on one route does not cover the others.

Part III: Platform and service compromise

A model that passed evaluation can still be exposed through acquired software, shared infrastructure, or the identities and interfaces of the running service.

Part IV: Untrusted instructions and excess authority

Untrusted text can influence model output, while document access, action authority, and execution remain separate security decisions.

  • Chapter 10: Indirect prompt injection
    Untrusted source text can influence model output. Protection depends on how the application interprets that output, what data it accepts, and which resulting actions and information transfers the receiving services permit.
  • Chapter 11: Retrieval access and disclosure
    Retrieval controls determine which records may reach the employee, model provider, and other recipients, including after permissions or stored copies change.
  • Chapter 12: Agent action authorization
    Action authorization keeps model-generated proposals separate from user, service, and policy decisions that permit consequential changes.
  • Chapter 13: Limits on agent execution
    Browser sessions and coding workspaces give agents access to files, credentials, and services, so execution limits and checks on each result must match the authorized task.

Part V: Security evaluation and response

Release tests measure behavior under specific conditions, while live monitoring and incident investigation check controls during use and recovery.

Part VI: AI in attack and defense

Claims about AI in cybersecurity require separate evidence for attacker assistance, defensive value, and deployment cost.

Part VII: Organizational AI governance

An organization needs to know which AI systems it uses, who can make decisions about them, and which risks justify further security work.

Part VIII: Security under changing conditions

Controls and evidence can lose validity when the sector, operating conditions, or system capabilities change.

Appendices

  • Implementation and research resources
    Tools, case collections, and research groups that support the book’s controls and evidence methods.
  • Conditions for robust aggregation
    The assumptions behind Krum, coordinate median, and trimmed mean results, for readers who need to check them before relying on these rules.
  • Glossary
    Definitions of terms used across the book, grouped alphabetically.
  • References
    Cited sources retain their reference numbers. Broader background is listed separately as further reading.