6  Data leakage and privacy attacks

An AI application can expose protected information through the data it sends, the copies it keeps, or the answers its model returns to specially chosen queries. Each route needs its own evidence, because a control on one route does not cover the others.

Suppose an employee uses a support assistant to draft a reply to a customer ticket. The ticket contains an account email, free text describing the problem, and an internal fraud flag. The application runs its own retrieval and then sends selected fields to an external model service, which returns drafted text. Protected information could leave the organization inside that provider request. Copies could also remain in logs, caches, or backups. If protected records were used for training, a separate query attack might recover memorized content or infer information about those records. Sending a prompt to a provider does not itself establish that the provider trains on it. The question is what each of these routes can disclose and what would show that it did.

Assume the organization has approved collection and processing for particular purposes and recipients. Those permissions set the rules for handling the data. They do not show that the application enforces them. The allocation of data rights and responsibilities is developed in Chapter 21 (Responsibility and legal duties). The technical distinction is between information carried directly by application requests or stored copies and information exposed through a model’s learned behavior. Training, inference, and retrieval create different paths, even when they use records from the same collection.

Protected data can leave through a provider request or remain in logs, histories, caches, and backups. A separate route lets an attacker query a model for memorized text, facts about individuals, or information about its behavior. The diagram separates these routes and the controls that apply to each. It also distinguishes measured attack results from a formal privacy bound when differential privacy is used.

Chapter 6 technical map has three panels. The first follows protected data through purpose and recipient checks, provider or internal service, stored copies, and access, retention, and deletion controls. The second follows attacker queries through a limited interface to observed responses, with distinct extraction, membership, attribute, and model-imitation targets plus empirical evidence. The third shows optional differentially private training and its accountant separately from controls on provider transfers, stored copies, and query outputs. A separate release-controller annotation assigns withholding to the controller while the privacy accountant tracks and computes privacy loss.
Figure 6.1: Data can leave through approved or unapproved transfers, remain in provider or organizational copies, or be learned through repeated model queries. The first branch follows purposes, recipients, fields, and retention. The second distinguishes extraction, membership and attribute inference, and imitation of model behavior under an observed interface. Empirical attack tests and a differential-privacy guarantee answer different questions and require separate evidence.

6.1 Data uses and handling rules

A ticket can be used to draft one answer, prepare a later training set, or investigate a service error. Each use creates different recipients and stored copies. Before evaluating a disclosure route, the organization needs a record of the permitted uses and a way to check where the data actually goes.

6.1.1 Data classification and purpose

Before approving a data use, the organization needs to know which records it requires, what processing will occur, and why. It also records which party will receive the result. Data classification groups records by the harm that disclosure, alteration, or loss could cause.

The NIST Privacy Framework1 separates the inventory of systems, people, data actions, purposes, and processing environments from the policies that govern them. Its Generative AI Profile2 also recommends checking training-data use for intellectual-property and privacy risk. These sources structure the decision but do not determine whether a specific use is lawful. A useful record identifies the data category and the processing purpose, which states the reason for a use. It also assigns a responsible internal role and lists every receiving party. That is the organization or service that obtains the data, including an external model provider. That record supplies the allowed-source and purpose rules that the technical checks below enforce.

Example

Classifying fields before transfer. Suppose the support ticket from the opening contains a ticket number, an account email, free text, and an internal fraud flag. For this constructed example, assume the free text has been checked and contains no fields excluded from the drafting purpose. The reference record permits that text and an internal ticket identifier for response drafting. The email remains inside the organization, and the fraud flag is excluded from this purpose. The internal ticket identifier is permitted so the application can match the draft to its case.

Assume context assembly keeps the ticket text, uses the permitted internal case identifier to associate the draft with its ticket, removes the email, and drops the fraud flag. The observed serialized request contains only the permitted fields. The conclusion is limited to this data class, purpose, recipient, and tested request path. It says nothing about whether the organization was entitled to process these fields in the first place. Authorization for the underlying data use belongs in the responsibility review in Chapter 21.

6.1.2 Data preparation and use permissions

An approved purpose for the system does not show that every training or evaluation record was collected with permission for that use. A source record is a data item kept with information about where it came from and the permission asserted for its use.

NIST Privacy Framework3 treats data actions, purposes, elements, and processing environments as inputs to privacy-risk management. NIST Generative AI Profile4 also recommends examining whether proprietary or sensitive training data are used consistently with applicable requirements. In common practice, a preparation job reads approved source records and filters disallowed fields. This supports data minimization, which removes fields or records that the stated purpose does not require. A documented rule controls duplicate removal before the job writes separate development and evaluation outputs. The evaluation output is an evaluation holdout, a test set kept outside parameter updates so it can measure behavior on unseen cases.

How much overlap between training and evaluation data distorts a score depends on the task and data, so the organization should treat overlap detection as an open test requirement. The organization also needs to retain the preparation decisions, so investigators can identify which output version contains a disputed record.

Under this preparation design, the job should enforce source and purpose rules before writing training or evaluation outputs. It checks the asserted permission, required fields, duplicate rule, and holdout assignment, then admits, changes, or quarantines a record. Job logs, rejected-record reports, and output digests show the decision. Coverage is reviewed records divided by all records in the prepared version, with overlap counts reported against both training and evaluation totals. The checks add processing and review cost. Incorrect source assertions or similarity outside the duplicate rule can bypass them. Recovery quarantines the version and rebuilds from the last accepted source set, without proving that the asserted permission was legally valid.

6.1.3 Data paths and copies

Wiz5 reported on September 18, 2023, that a shared access signature (SAS) token from Microsoft AI researchers exposed 38 terabytes of data. A SAS token is a signed link that grants access to cloud storage in Microsoft Azure. The token dated from 2020, about three years before the report. The 38 terabyte figure describes what the token made reachable, not what anyone copied. One access grant on shared storage can reach every copy kept there, whatever purpose each copy was made for.

Permission to use a source record for one purpose does not automatically cover every copy the system makes. A training corpus is the collection of examples used to update model parameters during training, while a prompt record affects one inference request. Numeric representations used for comparison or retrieval (embeddings) work as described in Chapter 5 (RAG and search poisoning).

The Privacy Framework6 calls for an inventory of data actions, data elements, purposes, and internal, cloud, or third-party processing environments. Applied to the reference system, the inventory follows source files into preparation jobs and their outputs. Separate entries cover indexes, embeddings, prompts, retrieved context, feedback, telemetry, and logs. Telemetry records operational events such as latency, errors, or resource use. This AI-specific list is a system decomposition based on common architecture rather than a structure prescribed by the framework.

A source document branches to a training corpus, prompt record, embedding index, telemetry and logs, and a provider copy across a recipient boundary. Each path has its own decision marker and register row for purpose, recipient, retention, deletion, and evidence. Training and inference uses remain distinct.
Figure 6.2: One source record can become training data, a prompt record, an embedding, operational telemetry, a log entry, or a provider copy. Each path requires its own permitted purpose, recipient, retained copy, deletion expectation, and evidence because approval for the source does not automatically cover later uses.

Derived forms should remain linked to their source and purpose because transformation does not by itself remove sensitive information. Each resulting path needs its own disclosure-control decision. Approval for the source does not automatically cover later uses. A check on one path says nothing about the copies sitting on another.

6.1.4 Data handling verification

Policies describe intended handling, but the review tests whether the running system follows them. NIST Privacy Framework Profiles7 compare current outcomes with target outcomes so an organization can identify gaps and plan improvements. NCSC guidance8 asks providers to explain where and how user data may be used, accessed, or stored. A statement about design or settings does not prove that the effective configuration or supplier behavior matches it.

Configuration evidence records effective settings and permissions, and contract evidence records supplier commitments. A review compares policy and contract with effective configuration, access logs, and controlled tests. A handling test submits controlled data and observes where it appears, who can retrieve it, and whether expiry changes that result. In one such test, a unique marker first enters a restricted document. The reviewer queries only approved paths before searching authorized logs and stores for that marker. The result is evidence about the tested path and time, not every future request. A passing review does not establish future supplier behavior.

A handling test also needs to follow data between components. An overbroad connector may ingest material the application should never see. A shared cache can return another user’s result, while an answer, link, or proposed tool call can carry sensitive text to a new destination. An end-to-end test marks synthetic sensitive text at the source. It then watches ingestion, retrieval, context, answer rendering, cache keys, network requests, and tool proposals. The aim is to locate the component that crossed the boundary instead of assigning every leak to the model. The connector, access service, cache, renderer, and action gateway each enforce a separate boundary using source, user, destination, and request state. If the observations are complete and trustworthy, a clean result shows that the monitored paths did not expose those markers during the test.

6.2 Sensitive data sent to providers

Bloomberg9 reported on May 2, 2023, that Samsung had restricted staff use of generative AI tools after employees pasted confidential material, reported as source code, into ChatGPT. The cited account is news coverage, rather than a Samsung incident report. It describes concern that the data might be disclosed to other users but does not establish that further disclosure. What left the company was ordinary text inside requests to an outside service.

With the permitted uses and copy inventory recorded, the first disclosure route to check is ordinary application traffic: fields the application sends outward and content it returns to the caller. The protection question is which fields may cross each boundary for the stated purpose, and the evidence is the data that actually crossed.

6.2.1 Disclosure controls

Once the paths are visible, the disclosure policy states which identity may send each field and identifies the permitted recipient. The Privacy Framework10 calls for managed permissions and authorizations based on least privilege and separation of duties. The UK National Cyber Security Centre (NCSC) and its partners11 recommend controls before potentially sensitive information is sent to an external application programming interface (API). These controls work at different points.

Permissions limit who can request data. A connector scope is the set of resources and operations granted to an integration, so it restricts what a service can fetch. Content checks apply to the permitted transmission. A data loss prevention check inspects content or metadata against disclosure rules. Redaction removes or masks selected content. Pseudonymization replaces direct identifiers while retaining a controlled way to reconnect records. Because detection quality depends on the data type and rule set, secret scanning, redaction, and pseudonymization each require tests against representative records. No universal detection rate follows from the guidance.

Example

Testing a seeded disclosure rule. Suppose the authorized task is a contract summary over retrieval followed by inference. The input contract contains one seeded test credential and a customer identifier. The employee may read the document, but the disclosure rule keeps both fields out of the request sent to the external model.

Assume the connector returns the document, the content check removes the credential, and pseudonymization replaces the customer identifier with customer-17. The observed serialized request lacks the credential and the direct identifier. The conclusion is that the tested rules changed this transfer. It does not establish their detection rate for other secret formats or identifiers. An encoded secret, a novel identifier, or a connector that sends data before the check can bypass the control.

The useful evidence here is the exact serialized provider request. The test reports blocked seeded items out of all seeded items and accepted permitted fields out of all permitted fields. It records added latency and manual-review volume alongside. More aggressive detection can reduce exposure while increasing false blocks and loss of useful context. Following the affected request identifiers into the copy register still does not show that every sensitive value is detectable.

6.3 Disclosure through stored copies

OpenAI12 reported on March 24, 2023, that a bug in redis-py showed some ChatGPT users other users’ chat titles on March 20. Redis-py is an open-source library that connects programs to a Redis cache. Partial payment details of about 1.2 percent of ChatGPT Plus subscribers active in a nine-hour window may also have been visible to others. These were name, email address, payment address, card type, the last four card digits, and expiry date. OpenAI said the actual exposure was “extremely low,” and the fault lay in the application’s library, not in the model. A cache is one of several stored copies that sit outside the main record and need their own rules.

The second route is storage rather than transmission: histories, caches, logs, backups, and secondary uses of collected data. A retention rule that mentions only the original file is incomplete, because each derived copy needs its own expiry and deletion evidence.

6.3.1 Retention and deletion evidence

Deleting a record raises two separate questions for each copy: when it must go, and what the deletion actually removes. A retention period specifies how long a copy may remain. Logical deletion removes ordinary access to a record, while media sanitization makes access to target data on storage media infeasible for a specified level of effort.

NIST13 defines media sanitization by the effort needed to recover target data and treats sanitization as a managed program. That definition is useful for storage media but does not establish deletion from a provider replica, conversation history, index, cache, backup, or learned parameter. The mechanism is a copy register in which each source record points to active stores, derived data, provider-held copies, backups, expiry rules, and the evidence available after deletion. Derived data are produced from source records, such as index entries, embeddings, or generated summaries. The review tests each registered location and reports where evidence is unavailable. A deletion request or interface response remains weaker than recovery testing, storage records, and provider evidence taken together. An unregistered replica or learned parameter lies outside this mechanism.

Example

Tracing deletion of one record. Suppose customer record C-104 reaches the end of its approved retention period. The copy register lists the source file, a search index entry, a cached passage, a conversation record, a daily backup, and a provider request identifier. The reference rule requires immediate logical deletion, cache invalidation within one hour, backup expiry after 30 days, and provider deletion evidence under the contract.

Assume the source, index, cache, and conversation checks return no accessible record. The backup remains until day 30, and the provider returns only a deletion receipt. For this example, the deployment owner is responsible for obtaining the contracted provider deletion evidence. The observed result supports removal from tested organization stores while leaving backup expiry and provider erasure as bounded claims. The deployment owner requests the missing evidence from the provider and records the unresolved claim about provider erasure in the copy register. Passing the organization-store checks does not prove physical erasure from an external system.

6.4 What query attackers observe

Application logs may contain more detail than a caller can see. A query-attack evaluation first identifies the interface outputs available to the attacker, then separates them from records available only to the evaluator. Query-only access, as introduced in Chapter 4 (Input attacks and response tampering), permits queries and observations without exposing model internals. The interface may return a label or generated text. Some systems also expose confidence scores or logits, the numerical scores before conversion to class probabilities. The attacker may also possess auxiliary data, information obtained separately from the target interface. A query budget limits how many requests the attack may make. In an adaptive query strategy, earlier answers guide the next requests.

NIST14 distinguishes black-box attackers by returned outputs, query access, and prior knowledge. These fields define a threat model but do not predict attack success. A rate limit bounds requests in a specified time interval, such as requests per minute. A total query budget instead bounds the number of requests across the whole test. The organization records the output fields, rate limits, account rules, adaptive-query allowance, prior knowledge, and target secret. Changing one can change the conclusion, as the following interface comparison illustrates. Allowed outputs can still disclose information, and distributed accounts or an unmetered route can bypass a budget. This fixed observation model supports the three query-based questions that follow: what was memorized, what can be inferred, and what can be imitated.

Example

Comparing what two interfaces expose. Suppose a hosted classifier is exposed through two interfaces for a privacy test on one candidate record. Interface A returns only a class label. Interface B returns the label and three class probabilities. Both allow 100 adaptive queries from one account. Assume both interfaces use the same fixed model. The evaluator submits the same 10 candidate inputs to each. They produce identical labels, while interface B also exposes how sharply the model separates the classes.

Interface B returns more information for the same queries: an analyst can use probability differences that interface A does not reveal. A disclosure result measured on interface B therefore cannot be assigned to interface A without another test.

6.5 Training data extraction

Nasr et al.15 reported on November 28, 2023, that asking the then-current ChatGPT model to repeat the word “poem” could make it stop repeating and emit memorized text. The authors report that, in the strongest configuration, over 5 percent of output was copied in sequences of at least 50 tokens. They checked outputs against an auxiliary collection of public web data, rather than the provider’s complete training corpus. This was a research test against the production service, disclosed to OpenAI before publication. The long matches support the authors’ memorization finding under that test, but do not identify every source record used by the provider.

The third route runs through the model’s parameters: content the model memorized during training and later reproduces for a querying caller. A fluent answer that resembles a training passage does not by itself establish training-data extraction. Investigators compare the returned content with the training corpus or another reliable record of what it contained. Establishing a confidentiality breach also requires checking who received the information and whether that recipient was authorized.

6.5.1 Training data extraction evidence

A claimed recovery of training content needs verification against the source records. Training-data extraction attempts to recover content from a model’s training data through its outputs.

Carlini et al.16 extracted verifiable examples from GPT-2 through a staged method. Candidate generation came first. A ranking step selected likely memorized text for checking against the training data. The experiments concern GPT-2 and candidate text checked against its training corpus, so the reported rates do not transfer to other models or sampling methods.

A canary is a controlled sequence inserted into training data to measure memorization. Earlier work by Carlini et al.17 defined exposure for inserted canaries as an empirical measure of unintended memorization. It measures how unusually easy a canary is to recover within the declared canary format and uniform randomness space. A canary value is sampled uniformly from a declared randomness space. Possible canaries are then ranked by model perplexity, a measure based on the probabilities the model assigns to their tokens. Lower perplexity places a candidate earlier in this ranking. Exposure is the binary logarithm of the space size minus the binary logarithm of the inserted canary’s rank. Exposure is relative to that format, space, model, and ranking rule. It is not normalized across different randomness spaces. A low result for one canary family does not prove that other memorized content is absent.

Example

Computing one canary’s exposure. Suppose an evaluator inserts one uniformly sampled canary from a format with 1,024 possible values, trains the model, and ranks all values by the trained model’s perplexity. Assume the inserted value ranks eighth, counting the most likely value as rank 1.

Its exposure is \(\log_2(1{,}024) - \log_2(8) = 10 - 3 = 7\) bits. The canary is eighth in this ordering of 1,024 candidates. This measures its unusual rank under the declared format. It is not seven bits of recovered personal information, a measured count of API queries, or a privacy guarantee for other records.

These results establish methods and examples, not a universal extraction rate. The extraction check records verified recoveries out of all generated candidates, while a canary test separately reports recovered canaries out of all inserted canaries with the query budget used. A prompt strategy or memorized sequence outside the chosen distribution can bypass the test, and candidate generation with restricted-data comparison can itself be costly. Restricting the interface and reviewing the data and training path are possible responses, but they cannot retract a verified sequence already disclosed. Differential privacy, discussed after the query attacks, can bound the influence of protected training records when the training mechanism and releases satisfy its assumptions.

6.6 Inference of private information

An attacker can also infer something private without reproducing a stored passage. The secret might be whether a record was used in training, what value a missing field holds, or what property a training population has. Reconstruction goes further by attempting to recover a record or a close representation. Each question needs its own target secret, auxiliary knowledge, correctness criterion, and baseline.

6.6.1 Membership inference

A person may want to keep even their participation in a sensitive training collection private. Membership inference estimates whether a record was used in training.

Shokri et al.18 introduced a membership-inference attack using prediction vectors and models trained by the attacker. Training fits a model to its training examples, so its losses or prediction patterns may differ between examples it trained on and otherwise comparable examples it did not see. The attack tries to learn that difference and relies on it transferring from the attacker’s training setting to the target. Overfitting can contribute, but it is neither the only possible source of the difference nor a prerequisite for every membership attack. High confidence by itself does not prove membership.

A shadow model imitates the target’s training setting using data whose membership the attacker knows. The attacker trains it on one set and queries it on both that set and a separate holdout. Each returned probability vector becomes an attack-training example labeled member or non-member according to that known split. These labels describe participation in training. The original task label, such as a disease class, describes what the model predicts and is a separate field.

The attacker pools these examples across shadow models and trains a membership classifier for each original task class. For a candidate of known class, the attacker obtains the target model’s full probability vector and submits it to the corresponding membership classifier. Shadow data supply the training labels that are unavailable for the target’s private records. The attacker knows the input and output formats. The shadow models use similarly distributed data and imitate the target model type, architecture, and training algorithm, or are trained through the same service. The paper does not establish label-only equivalence or universal success rates. NIST19 also summarizes loss-based and shadow-model attacks through a challenge game.

An evaluation samples members and non-members under a stated distribution. It reports the true-positive rate, the fraction of members correctly identified, and the false-positive rate, the fraction of non-members incorrectly called members. A high false-positive rate can make an apparent advantage unusable when real members are rare.

Matched member and nonmember samples share a target model, observation interface, attack score, and threshold. Known membership enters only the evaluator's confusion matrix. True positive rate uses all members as its denominator, and false positive rate uses all nonmembers. The test record states population, observation, query budget, and uncertainty.
Figure 6.3: A membership-inference test compares member and matched non-member cases under the same model interface and query budget. The threshold converts attack scores into predictions, so the reported operating point pairs true positive rate with false positive rate for the tested population.

Example

Reading a membership result. Suppose the operating mode is evaluation with a balanced set of 100 known training members and 100 non-members from the same reference population. An attack score and threshold classify 34 members as members and incorrectly classify 18 non-members as members.

The observed true-positive rate is 34 / 100 = 0.34, and the false-positive rate is 18 / 100 = 0.18. Among the 52 positive decisions, only 34 are correct. The conclusion is that the attack separates these test groups at this threshold, but the result does not establish usefulness when real members are rare or the target population differs. To see the effect of rare membership, suppose the same rates held among 100 members and 9,900 non-members. They would give 34 true positives and 1,782 false positives. Only 34 / 1,816, about 1.9 percent, of positive decisions would be correct. This calculation assumes the rates transfer to that population, which would need testing. A different population or attack strategy requires another evaluation, and released information cannot be taken back.

6.6.2 Reconstruction, attribute and property inference

Several privacy attacks are often grouped together even though they answer different questions. Data reconstruction seeks a record or a close representation, and a semantically similar reconstruction need not be the original record. Attribute inference estimates a missing field about an individual. Property inference estimates a feature of the training population, and a population property is not an individual attribute.

A model can reveal a missing input because its output depends on that input. Fredrikson et al.20 studied this route in a model predicting a patient’s stable warfarin dose from personal and genetic information. The attacker knew the patient’s actual stable dose, age group, race, height, and weight, or, in a stronger setting, every input except one genetic field. The interface returned a numerical dose prediction for submitted candidate records. The attacker also knew possible field values, population frequencies, and model-error statistics. Candidate records were weighted by their prior frequencies and by how well each predicted dose agreed with the known dose. Weights for the same missing genetic value were combined to select an estimate. The study compared estimates with recorded genetic values and with guesses from population frequencies without the target model.

Example

Inferring one missing field. Suppose a simplified test follows that stronger knowledge setting: the attacker knows all fields except a genetic category, represented by A, B, or C. For this numerical illustration only, the model predicts the known outcome exactly. Outcomes are normalized scores with no medical units, and all numbers below are invented.

The known outcome is 0.50. Holding the known fields fixed, queries with A, B, and C return 0.20, 0.50, and 0.80. Only B matches, so the attack predicts B. The evaluator’s withheld reference record also contains B. Without the target model, a baseline that always chooses the most frequent category selects A, assumed to occupy 60 percent of the reference population. That baseline is wrong on this case even though it would be correct on 60 percent of that population.

Conclusion. This trace shows one correct attribute inference and one incorrect baseline guess. It supplies neither an attack rate nor evidence of training membership. With model error or additional unknown fields, exact matching must give way to the weighted candidates described above.

The evidence required changes with the secret. Recovering an entire patient record would require comparison of all claimed fields with the original record under a specified matching rule. Plausible or semantically similar content is insufficient to establish exact recovery. A property attack instead asks, for example, whether more than half of the training population has category B. An illustrative test could train reference models using the target’s training procedure on populations with known proportions below and above one half. It would query each on the same probe records and test whether output patterns distinguish the groups on held-out reference models. If they do, the target model’s probe outputs support a group prediction, checked against its actual training proportion and a baseline using only prior population knowledge. Guessing one patient’s category cannot establish that proportion. NIST21 separates these targets and their access assumptions.

An evaluation reports correct inferences out of all eligible targets, with its reference answers, baselines, and false positives visible. Another secret, stronger auxiliary knowledge, or a shifted population can escape the test. The warfarin study establishes a method under its historical data and access conditions, not a general recovery result.

6.7 Model stealing

Carlini et al.22 reported partial recovery of production language models in March 2024. Their attack used interfaces returning the top five token log-probabilities, logarithms of the probabilities assigned to possible next tokens, with caller-controlled logit bias, an offset added to selected token scores before probability conversion. Bias made selected tokens appear among the returned entries. Repeated queries recovered relative token scores while accounting for the known bias and probability normalization.

The model first forms a vector of internal values. Its final projection layer multiplies that vector by a learned matrix to produce one logit per vocabulary token. The hidden size is the width of that internal vector. Across varied prompts, recovered score vectors reveal this size and a matrix spanning those directions, provided the prompts expose all internal directions and the projection preserves them. Changing internal coordinates together with the matrix leaves token scores unchanged, so the recovered layer need not reproduce the provider’s exact matrix. Ada’s recovered hidden size was 1,024 and Babbage’s was 2,048, with each recovery costing under $20 at the historical prices and interface settings. This recovered one layer and a size, not all weights. Providers subsequently changed their APIs.

Output restriction removes or coarsens information returned by the interface. Removing numerical outputs affects this recovery route; restricting caller-controlled bias changes the input controls instead. The paper also studies costlier variants without returned log-probabilities that retain bias control. These results do not establish extraction from an arbitrary text-only interface. A separate extraction goal trains a copy to imitate target responses. Its evidence is behavioral agreement on new inputs, which differs from recovering a layer up to equivalent coordinates.

6.7.1 Measuring model extraction

A model owner may need to protect the service’s learned behavior as well as its training records. Model extraction uses observations from a target model to build a copy or recover some of its parameters.

Tramèr et al.23 defined black-box model extraction through prediction interfaces and evaluated copies by agreement with the target over an input space. Functional agreement measures how often the two produce the same result over a stated input distribution. The general extraction goal is an equivalent or near-equivalent model, measured by agreement over a stated test distribution or a uniformly sampled input set. Agreement is behavioral evidence on the tested distribution: it does not by itself show identical weights or parameters. The paper also recovered exact parameters for some model classes. This worked when the interface exposed confidence values and the attacker could form the required equations from observed pairs. That stronger result should be reported separately from agreement, since it applies only under those model forms, feature handling, and sufficient queries. The online results are bounded to the studied model classes and services. There the authors obtained the model class and several settings from the service or its documentation, then tested guesses for the rest.

A useful test separates queries used to build the copy from unseen inputs used to measure agreement. It records query cost with the baseline agreement of a model trained without target outputs. The query gateway can reduce returned detail and enforce account budgets, blocking abusive access under the stated rule. Distributed identities or another interface can bypass the control, and output restriction may reduce legitimate use.

NIST24 discusses differential privacy for protecting training data and separately describes limiting or detecting suspicious model queries as possible mitigations for model extraction. It does not claim that these controls remove every disclosure path. NCSC guidance25 recommends controls on query interfaces to detect or prevent attempts to access confidential information, as recommended practice rather than a measured privacy guarantee. Recovery rotates credentials or changes the exposed interface and retests agreement. It cannot erase an extracted copy already built, and low agreement does not prove that the original weights stayed private.

6.8 Differential privacy

The preceding attacks ask different questions about training records and model behavior. Testing selected attacks cannot bound every possible use of released information. Differential privacy, or DP, bounds how much a randomized mechanism’s output distribution may change when one protected unit changes. A randomized mechanism is a computation whose random choices affect its output. The protected unit may be one record or one person’s full contribution. The neighboring-data relation specifies which two datasets count as differing in one protected unit. The parameters epsilon and delta quantify the allowed distributional difference under the formal definition.

Applied to training, this bound concerns the influence of protected records on released model outputs, whether an attacker seeks a memorized passage or an inference. It does not promise zero attack success or keep the released model’s behavior secret (Dwork and Roth26, 27).

Dwork and Roth28 define differential privacy over neighboring data sets and a specified randomized mechanism. Dwork and Roth’s composition results29 show that privacy loss accumulates across repeated mechanisms, and the applicable accounting method must match the mechanisms and sampling assumptions. The formal relation makes those dependencies explicit. Let \(D\) and \(D'\) be neighboring data sets under the declared adjacency relation. In Dwork and Roth’s default histogram model, they differ by one added or removed record. This is a person-level relation only when one person contributes at most one protected record, or when all of one person’s contributions are treated as the unit. For a randomized mechanism \(M\) and any possible set of outputs \(S\),

\[ \Pr[M(D) \in S] \le e^{\varepsilon}\Pr[M(D') \in S] + \delta. \]

Here \(\varepsilon\) limits the multiplicative change in output probability, while \(\delta\) is additive probability slack in this inequality for every neighboring pair and output event. Delta should not be read as a universal probability that the mechanism fails or that an individual loses all privacy. The probability is over the mechanism’s random choices. Here \(e\) is the base of natural logarithms, and \(\varepsilon\) and \(\delta\) are dimensionless. This \(\varepsilon\) is a privacy parameter, distinct from the input-change budget in Chapter 4. A related but different result, group privacy, scales a pure-DP guarantee with the number of changed records. Group privacy is not composition, which concerns repeated mechanisms (Dwork and Roth30).

The guarantee covers only outputs of the analyzed mechanism. Data disclosed by an unaccounted bypass path receive no differential-privacy protection from that mechanism. The bypass does not erase the mathematical guarantee for outputs that did pass through the mechanism, but it breaks any broader claim that the whole system protects all such data. Its meaning also depends on the protected unit and neighboring relation. Small privacy parameters do not establish secure storage, authorized recipients, or complete deletion. The guarantee also does not prevent learning facts that hold broadly across a population. Differential privacy does not encrypt, delete, or redact records.

The guarantee applies when the implemented mechanism satisfies the analyzed assumptions. A privacy accountant calculates a bound from the chosen mechanism, sampling rule, parameters, and releases. That calculation does not verify the implementation by itself. A release controller can use the account to withhold another release when a configured budget would be exceeded. Evidence therefore includes the actual implementation and configuration as well as the account. Epsilon and delta are guarantee parameters, not empirical attack-success fractions. Task quality is measured separately. Omitting a private release from the account can understate the combined bound. A raw-data bypass creates a separate unprotected disclosure, while the mathematical guarantee for the analyzed outputs remains as specified.

Stochastic gradient descent updates model parameters using error measurements from sampled training examples. A parameter gradient gives the local change in loss for each parameter, so a step in the opposite direction aims to reduce that loss. Unlike the input gradient used in Chapter 4’s evasion attack, this gradient changes the model rather than its input.

Differentially private stochastic gradient descent, usually shortened to DP-SGD, limits each example’s contribution before adding noise. It scales down any example gradient whose Euclidean length exceeds a clipping limit \(C\). The clipped gradients are summed. Independent zero-mean Gaussian noise with standard deviation \(\sigma C\) per coordinate is added before division by the specified lot size. Here \(\sigma\) is the dimensionless noise multiplier, \(C\) has gradient units, and the lot is the group of sampled examples contributing to one update. A training step then updates the parameters from this noisy average (Abadi et al.31).

Clipping supplies the connection to neighboring datasets. At fixed parameters, adding or removing one sampled record changes the clipped sum by a vector of length at most \(C\), provided other contributions and the sampling probability stay fixed. A replacement can change two contributions, so its bound can be \(2C\). Noise calibrated to the chosen relation makes these neighboring output distributions overlap in a controlled way. Arbitrary noise on an unbounded ordinary gradient supplies no such guarantee. This contribution bound concerns one record. Several records from one person require a bound and accountant for that person’s combined contribution.

Example

One noisy training update. Suppose a model has one adjustable parameter, currently \(\theta = 2.00\). A reference dataset has \(N = 100\) records. Each is sampled independently with probability \(q = 0.02\), and the specified averaging divisor is \(L = qN = 2\). The protected change is adding or removing one record, with \(q\) and divisor \(L\) held fixed across neighboring datasets. On this step, two records happen to be selected. Their loss gradients with respect to the parameter are 0.50 and 3.00. Set the clipping limit \(C = 1\), noise multiplier \(\sigma = 1\), and learning rate \(\eta = 0.10\).

The first gradient is already below \(C\) and stays 0.50. The second is scaled by \(1/3\) to 1.00. Their sum is 1.50. For this trace, assume the Gaussian noise draw is +0.20 from a distribution with mean zero and standard deviation \(\sigma C = 1\). The noisy average is \((1.50 + 0.20)/2 = 0.85\). The training update subtracts the learning rate times that average: \(\theta_{\text{new}} = 2.00 - 0.10(0.85) = 1.915\).

Removing the second sampled record changes the clipped sum by 1.00 rather than its original 3.00 contribution. The fixed divisor remains 2 even if a different number of records is sampled. The displayed noise value illustrates one draw. Repeating it deterministically would change the mechanism. This calculation shows how clipping, noise, averaging, and parameter updates fit together. It does not calculate a training privacy guarantee.

Abadi et al.32 introduced a practical accounting method for this training process. At each step, examples enter according to the sampling rule. The moments accountant tracks bounds describing the distribution of privacy loss, combines them across training steps, and converts the result to an \((\varepsilon, \delta)\) guarantee. The full derivation is in Section 3.2 and Appendices A and B of their paper. The proof assumes each example is included independently with probability \(q = L / N\). Here \(N\) is the reference number of training records, and \(L\) is the specified lot size, or expected sample count under that rule. The accountant combines this inclusion probability with the clipping and noise settings, number of steps, and outputs released. The bound is on the randomized process, not on whether an observed draw happens to be large. The paper’s implementation description instead permutes and partitions the data. The paper itself notes this difference, so the accountant must match the actual sampling rule.

Their reported privacy, compute, and accuracy results concern MNIST and CIFAR-10 image classification. The CIFAR experiment learned its convolutional layers from public CIFAR-100 data while training later layers privately. The paper names recurrent language modeling as future work, so it is not direct evidence about present-day large language model training. Per-example gradient work, the sampling rule, the noise setting, and the accountant can increase compute and reduce task quality. Whether rare records lose more quality requires testing on the actual model and data. Bounded per-example influence is also not evidence that poisoning or backdoors are prevented.

Example

A one-bit randomized response. Suppose the task is to release one person’s yes-or-no response with less direct disclosure. The input is the true bit. The mechanism reports the truth with probability 0.75 and the opposite with probability 0.25. If neighboring inputs change from yes to no, the largest probability ratio for either reported value is 0.75 / 0.25 = 3.

This mechanism satisfies the relation with \(\varepsilon = \ln(3) \approx 1.10\) and \(\delta = 0\) for that one bit. This is a distributional bound, not a promise that one output is false. The bound permits some learning about an individual.

Example

Composing two releases. Suppose the organization reruns this one-bit mechanism twice on the same protected input with fresh random draws and releases both results. Basic composition gives an upper bound of \(2\ln(3) \approx 2.20\) with \(\delta = 0\). The intermediate privacy record adds the two \(\varepsilon\) values.

The bound concerns two mechanism releases. Repeated predictions computed only from one already released DP model are post-processing. By themselves they do not spend the training privacy budget again, unless they trigger a new access to protected data or a new private release (Dwork and Roth33). Other mechanisms, dependence, and sampling can require a different accountant, so the example does not set a deployment budget.

6.9 Choosing privacy protections

The applicable defense depends on the disclosure path. Output restriction removes or coarsens returned information, and a rate limit bounds requests in a specified time interval. Both act on the interface path without altering memorized training data.

An empirical attack test measures selected attacks, while a formal guarantee follows from a defined mechanism and assumptions. Evidence should pair task quality with attack success, false-positive rates, query cost, privacy parameters where used, and operating cost, without collapsing unlike records into one score. New interfaces, identities, models, or data paths can bypass earlier evidence. These results provide bounded evidence for the stated disclosure mechanisms. Nothing here shows that ordinary fine-tuning establishes a privacy guarantee. Neither empirical non-disclosure on tested queries nor a formal bound under its assumptions covers data that left through another path.

An investigator who finds protected information in an answer needs to distinguish retrieved content, retained application copies, and learned model behavior before selecting a test. Release evaluation (Chapter 14) develops comparable measurements, while live observation (Chapter 15) connects a reported disclosure to the actual request and stored copies. Those records can themselves contain sensitive information and need restricted access. Recovery (Chapter 16) addresses affected copies, credentials, and services. It cannot make a recipient forget information already received. Access to process memory or model files bypasses query controls and belongs to platform isolation (Chapter 8).

Note

Chapter checkpoint. Suppose the support assistant from the opening uses an external model. Its source ticket contains an email address, free text, and an internal fraud flag, plus a seeded test credential. A hosted classifier in the same system returns labels and probabilities. A membership test reports 34 true positives and 18 false positives on 100 members and 100 non-members, and a separate private release uses the one-bit mechanism twice. What can the organization conclude?

Answer. The serialized provider request should contain only the permitted drafting fields. The email stays inside the organization, and the fraud flag is excluded from this purpose. The credential is removed before transfer, and each deletion is traced through the copy register. The membership test has a 0.34 true-positive rate and a 0.18 false-positive rate at the tested threshold, tied to its interface, query budget, and balanced population. Basic composition for the two private releases gives about 2.20 for \(\varepsilon\) with \(\delta\) equal to zero under the stated neighboring relation, covering only those mechanism outputs. Each result measures one stated property and leaves other disclosure routes, acceptable operating points, and unobserved copies unresolved.


  1. National Institute of Standards and Technology, NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, version 1.0 (2020), NIST CSWP 01162020, Appendix A, Table 2, ID.IM-P1-P8 and GV.PO-P1-P6, printed pp. 20 and 22, source. The inventory is separated from governing policies. These outcome categories do not form a complete AI data-classification scheme or legal permission test.↩︎

  2. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (2024), Section 3, action MP-4.1-010, printed p. 27 (PDF p. 31), source. A suggested diligence action. It does not decide whether a particular use is lawful.↩︎

  3. National Institute of Standards and Technology, NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, version 1.0 (2020), NIST CSWP 01162020, Appendix A, Table 2, ID.IM-P4-P7, printed p. 20, source.↩︎

  4. National Institute of Standards and Technology, Artificial Intelligence Risk Management Framework: Generative Artificial Intelligence Profile, NIST AI 600-1 (2024), Section 3, action MP-4.1-010, printed p. 27 (PDF p. 31), source.↩︎

  5. Hillai Ben-Sasson and Ronny Greenberg, “38TB of data accidentally exposed by Microsoft AI researchers,” Wiz blog (September 18, 2023), source. The researchers who found the exposure describe what the token made reachable, which is not a record of what others downloaded.↩︎

  6. National Institute of Standards and Technology, NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, version 1.0 (2020), NIST CSWP 01162020, Appendix A, Table 2, ID.IM-P4-P7, printed p. 20, source. The framework does not prescribe an AI-specific list of prompts, embeddings, feedback, caches, and logs. The system decomposition here follows common architecture.↩︎

  7. National Institute of Standards and Technology, NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, version 1.0 (2020), NIST CSWP 01162020, Privacy Framework Basics, “Profiles,” printed p. 8, source. A profile comparison structures review. It is not evidence that provider behavior matches a setting or contract.↩︎

  8. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), Secure deployment, “Make it easy for users to do the right things,” PDF p. 15, source. Transparency describes the stated design. It does not prove actual handling.↩︎

  9. Bloomberg News, “Samsung Bans ChatGPT, Google Bard, Other Generative AI Use by Staff After Leak,” Bloomberg (May 2, 2023), source. A news report with no Samsung incident report behind it, and the nature of the material is reported rather than confirmed.↩︎

  10. National Institute of Standards and Technology, NIST Privacy Framework: A Tool for Improving Privacy through Enterprise Risk Management, version 1.0 (2020), NIST CSWP 01162020, Appendix A, Table 2, PR.AC-P4, printed p. 26, source. The outcome does not prove that a particular connector scope or redaction rule is sufficient.↩︎

  11. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), Secure design, “Design your system for security as well as functionality and performance,” PDF p. 10, source. Design guidance. It does not evaluate any specific data-loss-prevention product.↩︎

  12. OpenAI, “March 20 ChatGPT outage: Here’s what happened,” OpenAI (March 24, 2023), source. The company’s own account of a bug in an application library, including its own estimate that actual exposure was extremely low.↩︎

  13. Ramaswamy Chandramouli and Eric A. Hibbard, Guidelines for Media Sanitization, NIST SP 800-88 Rev. 2 (2025), abstract and Section 2, with the sanitization program in Section 4, source. The definition concerns media and does not establish deletion from every logical or provider-managed copy.↩︎

  14. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, National Institute of Standards and Technology (2025), Section 2.1.4, report pp. 8-9, and Section 2.4, report pp. 28-33, source. The categories define a threat model. They do not predict attack success.↩︎

  15. Milad Nasr, Nicholas Carlini, Jonathan Hayase, et al., “Scalable Extraction of Training Data from (Production) Language Models,” arXiv:2311.17035 (November 28, 2023), paper, author explanation. See Sections 4-5 of the author manuscript and the linked author explanation. The rate applies to the strongest tested configuration. For the closed model, the authors matched long output sequences against an auxiliary public-data collection, not a complete copy of the provider’s training corpus. The paper records disclosure to OpenAI on August 30, 2023.↩︎

  16. Nicholas Carlini et al., “Extracting Training Data from Large Language Models,” Thirtieth USENIX Security Symposium (2021), pp. 2633-2650, Sections 3-6, especially candidate generation and ranking in Sections 4-5 and corpus verification in Section 6.1, source. The experiments concern GPT-2 and candidate text checked against its training corpus.↩︎

  17. Nicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos, and Dawn Song, “The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks,” Twenty-eighth USENIX Security Symposium (2019), pp. 267-284, Section 4, especially the exposure definition in Section 4.2, printed pp. 270-272, source. Exposure is defined relative to the canary format, space, model, and ranking rule, not every possible secret.↩︎

  18. Reza Shokri, Marco Stronati, Congzheng Song, and Vitaly Shmatikov, “Membership Inference Attacks Against Machine Learning Models,” IEEE Symposium on Security and Privacy (2017), pp. 3-18, DOI 10.1109/SP.2017.41, Sections IV-V, especially Section V-D and Figure 3 on known shadow membership and per-class attack training, source. The method uses full prediction vectors, true labels, and similarly trained shadow models. Its experiments do not establish a universal success rate or label-only equivalence.↩︎

  19. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, National Institute of Standards and Technology (2025), Section 2.4.2, report pp. 29-30, source. A taxonomy and literature summary, not a deployment-specific result.↩︎

  20. Matthew Fredrikson, Eric Lantz, Somesh Jha, Simon Lin, David Page, and Thomas Ristenpart, “Privacy in Pharmacogenetics: An End-to-End Case Study of Personalized Warfarin Dosing,” 23rd USENIX Security Symposium (2014), pp. 17-32, Sections 3.1-3.2 and Figure 2, printed pp. 20-23 (PDF pp. 5-8), conference paper. The attack uses known dose and background fields, black-box dose predictions, population frequencies, and model-error information. Its historical study does not establish general reconstruction success.↩︎

  21. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, National Institute of Standards and Technology (2025), Sections 2.4.1 and 2.4.3, report pp. 28-31, source. A similar reconstruction is not the original record, and a population property is not an individual attribute.↩︎

  22. Nicholas Carlini, Daniel Paleka, Krishnamurthy Dj Dvijotham, et al., “Stealing Part of a Production Language Model,” arXiv:2403.06634 (March 11, 2024), source. Author version 1, Sections 3-5, 6, and 7.2 with Table 4. The production attack used top-five log-probabilities and controllable logit bias, recovering the final projection layer up to equivalent internal coordinates and the hidden size. The paper reports provider interface changes after disclosure.↩︎

  23. Florian Tramèr, Fan Zhang, Ari Juels, Michael K. Reiter, and Thomas Ristenpart, “Stealing Machine Learning Models via Prediction APIs,” Twenty-fifth USENIX Security Symposium (2016), pp. 601-618, Section 2 and Sections 4-6, source. The attacks study particular model classes and prediction services. Agreement alone does not establish identical parameters.↩︎

  24. Apostol Vassilev et al., Adversarial Machine Learning: A Taxonomy and Terminology of Attacks and Mitigations, NIST AI 100-2e2025, National Institute of Standards and Technology (2025), Section 2.4.5, report pp. 32-33, source. Listing a mitigation does not establish effectiveness for a particular attack or model.↩︎

  25. UK National Cyber Security Centre, CISA, and international partners, Guidelines for Secure AI System Development, version 1.0 (2023), Secure deployment, “Protect your model continuously,” PDF p. 14, source. Recommended practice, not a measured privacy guarantee.↩︎

  26. Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy, Foundations and Trends in Theoretical Computer Science 9(3-4) (2014), pp. 211-407, DOI 10.1561/0400000042, Definition 2.4, author manuscript pp. 17-18, source. The default histogram relation is one added or removed record. Person-level protection needs one record per person or an explicit person adjacency and contribution bound.↩︎

  27. Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy (2014), Proposition 2.1, “Post-Processing,” printed p. 19, and Theorem 2.2, pure differential privacy for groups, printed p. 20, author manuscript. Relevant passages read September 16, 2026. Post-processing concerns functions of the private mechanism’s output without a new access to protected data. Group privacy changes the protected relation and is distinct from composition of releases.↩︎

  28. Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy, Foundations and Trends in Theoretical Computer Science 9(3-4) (2014), pp. 211-407, DOI 10.1561/0400000042, Definition 2.4, author manuscript pp. 17-18, source. The default histogram relation is one added or removed record. Person-level protection needs one record per person or an explicit person adjacency and contribution bound.↩︎

  29. Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy, Foundations and Trends in Theoretical Computer Science 9(3-4) (2014), pp. 211-407, DOI 10.1561/0400000042, Section 3.5, especially Theorem 3.16, author manuscript p. 43, source. Basic composition sums per-mechanism guarantees. The accountant must match the mechanisms and sampling assumptions.↩︎

  30. Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy (2014), Proposition 2.1, “Post-Processing,” printed p. 19, and Theorem 2.2, pure differential privacy for groups, printed p. 20, author manuscript. Relevant passages read September 16, 2026. Post-processing concerns functions of the private mechanism’s output without a new access to protected data. Group privacy changes the protected relation and is distinct from composition of releases.↩︎

  31. Martín Abadi et al., “Deep Learning with Differential Privacy,” ACM CCS (2016), pp. 308-318, DOI 10.1145/2976749.2978318, Algorithm 1 and Section 3.1, PDF p. 3, Section 3.2, PDF pp. 4-5, and Appendices A and B, PDF pp. 12-14, author manuscript v2, October 24, 2016. The reported privacy, compute, and accuracy results concern MNIST and CIFAR-10 experiments, with public CIFAR-100 pretraining for the CIFAR convolutional layers, not present-day large language model training.↩︎

  32. Martín Abadi et al., “Deep Learning with Differential Privacy,” ACM CCS (2016), pp. 308-318, DOI 10.1145/2976749.2978318, Algorithm 1 and Section 3.1, PDF p. 3, Section 3.2, PDF pp. 4-5, and Appendices A and B, PDF pp. 12-14, author manuscript v2, October 24, 2016. The reported privacy, compute, and accuracy results concern MNIST and CIFAR-10 experiments, with public CIFAR-100 pretraining for the CIFAR convolutional layers, not present-day large language model training.↩︎

  33. Cynthia Dwork and Aaron Roth, The Algorithmic Foundations of Differential Privacy (2014), Proposition 2.1, “Post-Processing,” printed p. 19, and Theorem 2.2, pure differential privacy for groups, printed p. 20, author manuscript. Relevant passages read September 16, 2026. Post-processing concerns functions of the private mechanism’s output without a new access to protected data. Group privacy changes the protected relation and is distinct from composition of releases.↩︎