Glossary
This glossary gives short definitions for reference. The chapters explain how each method works, with examples and limitations.
Ablation: An experiment that removes or replaces one component to compare its contribution under otherwise shared conditions. Interactions with other components can affect the result.
Access token: A credential that grants limited access to one resource server.
Access-control list (ACL): A list of users or roles allowed to access a resource. The application checks it before protected information enters a model request.
Accuracy: The share of all cases classified correctly. It hides error types when one class dominates.
Action: A model or agent-policy output that proposes the next step, such as a tool call, clarification question, delegated subtask, or final answer.
Agent: A system in which a model selects actions toward a goal from observations, tools, and state, while application code enforces authorization, execution, budgets, and stopping rules.
Agentic RAG: A retrieval-augmented system in which an agent chooses and iterates retrieval actions from observations, within a bounded recovery loop.
Answer relevance: Whether a response resolves the requested information rather than only discussing the same topic.
Application Programming Interface (API): A defined interface through which application code sends a request to a provider and reads the response.
Application software: The code that forms requests from user input and company data, invokes the model, checks whether its answer is usable, and authorizes any next action.
Atomic update: A check and change treated as one indivisible operation, so concurrent requests cannot both act on an unchecked old value.
Attention: The transformer operation that lets each token representation combine information from other positions in the sequence.
Authorization: The application decision that a specific identity may perform a specific action on a specific resource under the current policy.
Authorization server: The system that verifies the user and issues a limited access token.
Base model: A pre-trained language model that primarily continues text. Alignment training adapts the model to follow instructions and conversational roles.
Benchmark saturation: A condition in which many models score near a benchmark’s maximum, leaving little separation between their scores. It differs from contamination: models can solve the items without having seen them in training.
BLEU: A machine-translation metric that combines modified n-gram precision with a brevity penalty.
BM25: A lexical ranking function that rewards matching query terms, weights rare terms more, and accounts for term frequency and document length.
Bootstrap confidence interval: An interval estimated by resampling evaluation cases with replacement and recomputing a statistic. The percentile method used in Chapter 4 selects quantiles of those resampled statistics. Its intended repeated-evaluation coverage depends on the sampling assumptions and interval method. It is not the probability that a particular candidate improves the product. See Uncertainty, release thresholds, and decisions in Uncertainty, release thresholds, and decisions.
Byte-pair encoding (BPE): A subword method that starts with small units and repeatedly merges the pair that occurs most often in the training text.
Canary release: A partial and time-limited deployment of a change, evaluated before full rollout.
Capture: Identifying candidate information from a conversation and associating it with its source.
Chain-of-thought prompting: Prompting a model to generate intermediate steps as text before its final answer.
Chat model: An instruction-following model used with a role-aware serialization that distinguishes system, user, assistant, and sometimes tool messages.
Chat template: The model-specific rule that converts system, user, and assistant messages into the control tokens and text the model was trained on.
Checkpoint: A saved model version containing a particular set of parameters and its associated configuration. Loading another checkpoint changes the model used for inference. A workflow checkpoint is a different object: saved execution state for one thread or session. See Session state and persistent memory in Session state and persistent memory.
Chunk: The source unit placed in a retrieval index and later returned as evidence.
Citation support: Whether the evidence attached to each claim supports that claim.
Client: In MCP, the host’s protocol component that connects to one server and exchanges messages with it. A host creates a separate client instance for each server it uses.
CodeAct: An agent pattern in which the model’s action is an executable program, run in an isolated sandbox.
Cohen’s kappa: An agreement measure that corrects observed agreement for the agreement expected by chance from each rater’s label rates.
Confidence calibration: The agreement between a model’s stated probabilities and its observed accuracy on labeled cases.
Confused deputy: A security failure in which software with legitimate authority is tricked into using that authority for an unapproved caller or purpose. See Trust at the protocol boundary in Trust at the protocol boundary.
Confusion matrix: A table that counts each actual class against each predicted class.
Constrained decoding: Restricting next token choices to continuations compatible with permitted syntax. Section 2.4 explains grammar-based constraints and their limits.
Context engineering: Selecting, ordering, labeling, trimming, and recording all information supplied to a model, not only the instruction wording.
Context relevance: Whether the retrieved passages are useful for the question.
Context selection: Choosing which policy, history, and retrieved data enter a request.
Context window: The model’s token capacity for the serialized request and, for APIs with a shared limit, the generated output.
Conversation summary: A compressed representation of earlier turns. It should preserve decisions, constraints, entities, and unresolved questions while discarding conversational filler.
Corpus: The authorized collection of documents or records available to the retrieval system.
Correctness: Whether an answer matches a trusted result or expert judgment.
Cosine similarity: A similarity measure based on the angle between two vectors. It equals the dot product when both vectors have length one.
Criteria drift: The change in evaluation criteria that occurs as people grade real outputs and refine what they expect.
Data contamination: A condition in which benchmark items or close variants entered model training or evaluation-time retrieval, so a high score may not show generalization.
Decoding strategy: The rule that turns next-token probabilities into a selected token, such as taking the most probable token or sampling.
Direct Preference Optimization (DPO): An alignment method that trains directly on preferred and rejected response pairs, without a separate reward model or reinforcement-learning step.
Document: In common RAG libraries, a text payload paired with metadata. Depending on loader configurations, it represents complete files, single pages, sections, or individual database records.
Edge: In a workflow graph, a connection specifying which node can run next.
Embedding: A learned vector representation associated with a token or another discrete object. Similarity in embedding space is useful but does not guarantee identical meaning or factual equivalence.
Embedding model: A model that maps a text unit to a fixed-dimensional vector used for similarity search.
Episodic memory: Timestamped records of earlier interactions, actions, and outcomes that an application may retrieve while access and retention rules permit.
Evaluation: A repeatable procedure that runs test inputs through a specific system version and scores the results to support release decisions.
Evaluation run (eval): A combination of fixed cases, a scoring procedure, a comparison target, and the product decision it informs.
Evaluation-driven development (EDD): The book’s workflow for treating each proposed change as an experiment with fixed inputs, explicit scoring, uncertainty, and a decision.
Evaluator-optimizer: A bounded loop that generates a candidate, grades it with an external evaluator, and revises it until the criteria pass or a budget expires.
Evidence hit at k (Hit@k): A retrieval measure that equals one when at least one labeled evidence item appears among the first k results for a question, and zero otherwise. See Retrieval metrics in Retrieval evidence metrics.
Evidence validation: A check that each field in a response follows from the supplied data and application rules.
F1: The harmonic mean of precision and recall.
Faithfulness: Whether an answer’s claims are supported by the context supplied to the model.
Few-shot prompting: Prompting with instructions plus worked examples inside the request.
Foundation model: A model trained on broad data before it is adapted or used for a particular product.
Function calling: A model interface that returns a structured proposal for a tool name and arguments. The application validates the proposal and executes only permitted calls.
Golden set: A versioned collection of evaluation tasks, each with an expected outcome or grading criterion.
Guardrail metric: A measured requirement that a proposed change must satisfy, such as a latency, cost, or unsafe-answer limit. It differs from a safety control that blocks an action.
Held out: Excluded from demonstrations and prompt tuning so a test case can assess behavior on an unused input. Section 2.3 applies this separation to prompt comparison.
Human-in-the-loop gate: A workflow pause for a person’s decision. Approval for an action is tied to its displayed parameters and relevant current state.
Hybrid retrieval: Retrieval that combines lexical and semantic candidate lists to catch both exact terms and paraphrases.
Idempotency key: An identifier used by a service to recognize repeated submissions of one operation rather than treat each retry as a new operation.
Idempotent operation: An operation whose intended effect is the same whether an identical request is applied once or several times. See Workflow state and transitions in Workflow state and transitions.
In-context learning: Using instructions or examples in the prompt sent for a task to influence the model’s response. The model’s stored weights stay the same.
Index: A data structure that maps query representations to candidate source units. It may be lexical, vector-based, structured, or hybrid.
Inference: The operating mode in which a trained model receives a request and generates output while its parameters stay fixed.
Instruct model: A base model further trained to respond to instructions, usually through demonstrations and preference-based alignment.
Join: A workflow step that waits for required branch results and applies a defined rule for combining them.
JSON-RPC: A JSON message format for requests, responses, and notifications that need no response.
Large language model (LLM): A foundation model that reads and writes text tokens.
LLM-as-a-judge: An evaluator in which a language model scores or compares candidate outputs against a rubric.
LLMCompiler: A planner that represents function calls as a dependency graph and dispatches ready calls in parallel.
Log-probability: The natural logarithm of the probability a model assigned to a token. Some APIs return it for each generated token.
Logit: The unnormalized score a model computes for one vocabulary token at a generation step. It can be positive or negative.
Long-term memory: Records an application saves and supplies in later requests, such as a user’s preferences or earlier decisions. Saving these records does not update the model’s learned parameters. See Selecting information for each request in Selecting information for each request.
Lost-in-the-middle behavior: The tendency of some models to use evidence near the start or end of a long context more reliably than equally relevant evidence in the middle.
Maximal Marginal Relevance (MMR): A selection method that trades relevance against redundancy with results already selected.
Mean Reciprocal Rank (MRR): The average over questions of one divided by the rank of the first relevant result, with zero when none is retrieved.
Memory system: The database, index, and application logic that manages persistent state across agent sessions.
Metadata: Source identity and attributes stored with a chunk, such as document, page, section, date, company, permission, or content type.
Mid-response steering: Additional user input received while a response is still in progress. The application or a supporting provider continues from updated conversation state. Earlier output and completed actions remain unchanged unless the application handles them separately. See Request, response, and validation cycle in Request, response, and validation cycle.
Model alignment: The process of making a model’s behavior better match intended instructions, preferences, or safety criteria in specified situations. Chapter 1 focuses on further training that adjusts the model’s parameters toward those behaviors. See Model alignment in Model alignment.
Model Context Protocol (MCP): A client-server protocol enabling an AI host to discover and invoke server-provided tools, resources, and prompt templates.
Multi-agent system: A system that coordinates two or more agents through explicit communication, routing, termination, and rules for which component may read or change shared state.
N-gram: A run of n neighboring tokens, used by overlap metrics such as BLEU and ROUGE.
Node: In a workflow graph, a function that reads state and returns updates.
OAuth: An authorization protocol through which a user allows a client limited access to a service without sharing the user’s password.
Observation: The environment result returned after an action, such as tool output, error, updated state, or user response.
Observed agreement: The fraction of audited cases for which two label sources, such as a human and a model judge, give the same label.
Online controlled experiment: Random assignment of users or sessions to two system versions, followed by a comparison of product outcome metrics. Also called an A/B test.
Orchestrator-worker pattern: A pattern in which an orchestrator decomposes a request at run time, assigns subtasks to workers, and combines their results.
Output specification: The fields, types, allowed values, evidence requirements, and missing-evidence response that make a model answer usable by the next program step.
p95 latency: The ninety-fifth percentile of measured response times, calculated with a specified quantile method. Ties and finite samples can change how many observations fall above it. See Uncertainty, release thresholds, and decisions in Uncertainty, release thresholds, and decisions.
Parallelization: A workflow pattern that runs independent subtasks concurrently and joins their results.
Parameter: A number learned during training and stored in the model. Parameters, also called weights, stay fixed during ordinary inference. Training updates them, while loading another saved model version, called a checkpoint, replaces the set used for inference.
Pass@k: A code-generation metric for the probability that at least one of k generated attempts passes the tests.
PKCE: Proof Key for Code Exchange, which binds an authorization code to a secret verifier held by the requesting client.
Plan-and-Act: A method that separates a high-level planner from an executor that carries out environment actions. Its replanning variant revises the plan from observations.
Plan-and-Execute: A planner that creates a full plan before acting and then runs its steps.
Planner: The model or software component that applies a planning policy to select actions, order dependencies, and propose a stopping condition. Application code validates execution and termination.
Positive class: The outcome a classification task treats as present, such as
requires escalation. It fixes which errors count as false positives and false negatives.Precision: The share of positive predictions that are correct.
Prefix: The token sequence already present when the model computes scores for the next position.
Procedural memory: Versioned prompts, tools, workflows, examples, or model adaptations that shape how a system performs a task.
Product requirement: A verifiable statement of required or forbidden system behavior under stated conditions, traced to a product outcome.
Promotion: The decision to make a captured candidate available as a reusable fact or procedure after checking supporting records, identity, intended use, confidence, expiry, and required authority. Human review is required only where policy calls for it.
Prompt: The input text that describes a task and can include instructions, examples, or both.
Prompt chaining: A workflow pattern that splits a task into steps whose intermediate outputs can be typed, inspected, and retried.
Prompt injection: An attack in which instructions placed in text the application treats as data change model behavior or tool use. It is direct when the user writes them and indirect when they arrive in retrieved or external content.
Provenance: The record of the supporting sources used to validate a stored fact.
Proxy metric: An offline measurement used in place of a product outcome observable only in production.
RAG: Retrieval-augmented generation: retrieving external evidence at request time and placing it in the model’s context so the answer can be grounded in it.
Ragas: A software library providing metrics for evaluating retrieval-augmented answers and their contexts. Metric results depend on the evaluator and assessment procedure.
ReAct: An agent-control pattern that alternates a model decision, one validated action, and the resulting observation before the next decision.
Reasoning: Using available information through one or more intermediate steps to reach a conclusion. In an LLM system, this may refer to internal model computation, generated intermediate steps, or application-managed stages. See Model alignment in Model alignment.
Reasoning tokens: Tokens a model generates before its final answer that a provider may hide, summarize, or count separately.
Recall: The share of actual positive cases that the system finds.
Reciprocal rank fusion (RRF): A fusion method that combines ranked lists by summing 1/(k + rank) for each candidate, ignoring raw scores on different scales.
Reducer: A rule for applying an update to a state field, such as appending a trace entry rather than replacing the trace.
Reflection: A revision pattern in which a model critiques its own or another model’s output before another attempt.
Reflexion: A feedback method that generates a verbal reflection from an evaluator’s feedback and the prior attempt, then saves it for the next attempt without updating model weights.
Reinforcement learning from human feedback (RLHF): An alignment setup in which response comparisons train a reward model, and reinforcement learning then makes higher-scoring responses more likely.
Release decision record: A versioned document that links the task definition, evaluation results, component traces, thresholds, blocking failures, and follow-up to a PASS or HOLD decision for specified conditions.
Request trace: A record of the policy, history, evidence, tools, token allocations, and other inputs actually sent for one model request.
Requirement traceability: The links from each requirement to its measurement, threshold, evaluation cases, and results.
Reranker: A model or scoring step that examines a query together with each retrieved candidate and reorders the results. Task evaluation determines whether the new order improves retrieval. See Staged retrieval pipeline in Staged retrieval pipeline.
Resource server: The server, such as an MCP server, that checks an access token before serving the requested capability.
Reward model: A learned scorer that assigns a score to a candidate response in its prompt context.
ReWOO: A planning method that separates planning from tool observations. Workers fill planned output variables with tool results, and a solver uses the plan and evidence to answer.
ROUGE: A family of n-gram overlap measures developed for summary evaluation.
Routing: A workflow pattern that classifies a request into a small explicit set and sends it to a specialized workflow.
Safety guardrail: A policy or validation step that restricts inputs, outputs, tool use, or state changes. It may use deterministic code, a model, or human review. Authorization for actions with external effects should not depend on model judgment alone.
Security boundary: A division enforced by software controls such as identity, permission, validation, or isolation. Prompt delimiters do not create one.
Selective retrieval: A history or knowledge policy that searches stored records and includes only items relevant to the current request.
Semantic memory: Curated facts or preferences stored separately from the events from which they were derived.
Semantic search: Retrieval that ranks evidence by embedding similarity instead of exact word overlap.
Service level objective (SLO): A target value or range for a measured service indicator, such as p95 latency or error rate.
Shadow deployment: Running a candidate on copies of live requests while users receive only the current system’s answers.
Slice: A subset of evaluation cases grouped by a relevant property, such as language, customer tier, risk level, query type, document length, or tool requirement.
Sliding window: A history policy that keeps the most recent turns and drops older ones. It is cheap and predictable but can lose earlier commitments.
Softmax: The function that turns logits into probabilities by exponentiating each score and dividing by the sum, so the results add up to one.
State machine: A process represented by its current state and allowed transitions to another state.
Streaming: Delivery of partial response events while generation continues. A streamed fragment is provisional until the response finishes and the application validates it. See Request, response, and validation cycle in Request, response, and validation cycle.
Structural validation: A check that a response has the required shape, fields, and types.
Structured output: An approach that declares the fields and value types receiving code expects. Supported schema constraints can restrict generation to permitted fields and values.
Subword token: A token that can be a whole common string or a fragment of a rarer word, so the same text can produce different token counts under different tokenizers.
Supervised Fine-Tuning (SFT): Further training on prompts paired with suitable responses, which makes those responses more likely.
System instruction: Application rules supplied in the interface’s designated instruction role. Their priority relative to user messages depends on the model and interface. These rules describe limits and required handling that persist across requests.
Task DAG: A directed acyclic graph whose nodes are executable tasks and whose edges state that one task depends on another task’s output.
Task specification: A statement of the goal, supplied data, allowed evidence, response fields, and missing-evidence behavior for a model task.
Temperature: For positive values, a generation setting that divides logits before softmax. Lower positive values concentrate probability on the highest-scoring tokens. Zero-temperature behavior is a separate API convention.
Token: A piece of text, such as a word, word fragment, punctuation mark, or byte, that a tokenizer maps to an identifier the model processes.
Tokenizer: The component that splits text into tokens and gives each token an identifying number from a fixed vocabulary.
Tombstone: A small record stating that an item was deleted or replaced, without retaining the removed value. Replicas and indexes use it to avoid restoring an old copy. See Semantic memory and correction in Semantic memory and correction.
Tool: An operation the software can execute, with a name, description, defined input types, and a result returned to the model or workflow.
Tool allowlist: The set of tools and tool schemas that application policy permits for the current identity and task.
Tool specification: A description of a capability’s purpose, inputs, outputs, errors, side effects, and required permission.
Top-k: A sampling rule that keeps only the k most probable tokens before sampling.
Top-p: A sampling rule that keeps the smallest set of most probable tokens whose total probability reaches a threshold.
Training: The operating mode in which software compares predictions with example targets or preferences and updates the model’s parameters.
Trajectory: The ordered record of model decisions, tool calls, observations, and state changes in one trial.
Transformer: A neural-network design in which attention lets each token representation combine information from other positions.
Transition: In a workflow, movement from one state or step to another under a defined rule.
Trial: One attempt at one evaluation task. Because sampling, retrieval, and tools can vary, the same system can produce different answers across trials.
Unigram: A subword tokenizer that scores possible splits of a word and selects a likely one.
User or resource owner: In OAuth-based MCP authorization, the person who grants permission.
User request: The current task or question. It can include user data, but the application must still treat that data as input rather than as a rule that overrides its own policy.
Veto requirement: A mandatory product requirement whose failure blocks release and produces a HOLD decision. See Product requirements and evaluation criteria in Product requirements and evaluation criteria.
Vocabulary: The fixed list of token identifiers a tokenizer and model share. Its entries may represent complete words, word fragments, punctuation, bytes, or other subword units.
Wilson score interval: An interval estimating a binomial success probability from a success count and trial count under independent, equal-probability trials.
WordPiece: A subword tokenizer that takes the longest matching vocabulary pieces from left to right.
Workflow: A predefined sequence or graph of model calls, tool calls, validations, and state transitions whose control flow is governed by application code.
Workflow state: The data available at a workflow step, such as an order ID, selected policy, draft reply, and status. It can be temporary or saved between calls.
Working memory: Active task state, such as the goal, plan and intermediate results. The application selects part of this state for the current model request.
Zero-shot prompting: Prompting with instructions but no demonstrations.