Glossary
MLOps combines vocabulary from infrastructure, distributed systems, model evaluation, storage, retrieval, and orchestration. The same term can connect several components, so each entry identifies the object, behavior, or measurement involved. The entries use consistent meanings across training, serving, evaluation, storage, and pipelines.
access-control list (ACL): A record of which users or groups may perform specified actions on a resource. A stored document ACL helps restrict retrieval, but the service must check current permissions when answering.
Access pattern: The measurable read/write, size, concurrency, sharing, durability, latency, and lifecycle behavior of data.
Activation: An intermediate layer output retained as needed for backpropagation during a training step.
Activation recomputation: Discarding selected forward-pass activations and recomputing them during the backward pass to reduce stored activation memory. Full recomputation costs roughly one extra forward pass. Selective schemes have different costs. Also called activation checkpointing.
Adapter (LoRA): A small set of trained low-rank matrices added to frozen base weights, so a fine-tuned variant can be stored and served without a full copy of the model.
Agent: An application in which a model repeatedly chooses the next action, such as requesting a tool. The application checks the request, may execute it, and returns the result to the model until it produces a final answer.
Agent trajectory: The recorded sequence of model, tool, observation, and state transitions made while an agent completes one task.
All-reduce: A collective that combines values from all ranks and returns the result to every rank.
API gateway: The request-entry component that checks identity and request shape, enforces quotas and policy filters, and records billing identity.
Approximate Nearest Neighbor (ANN) search: Retrieval that approximates exact nearest-neighbor results to reduce latency or storage cost. Some methods prune candidates. A flat compressed-code search can still scan every code.1
ANN recall@k: For positive k, the fraction of an exact search’s k reference IDs returned by an approximate top-k search using the same stored vectors, distance metric, filters, and declared deterministic tie rule. The benchmark requires at least k eligible vectors after filtering. Other queries are reported separately or use a predeclared smaller cutoff. It measures approximation fidelity, not human relevance.
Artifact: A versionable product of a run, such as a checkpoint, dataset manifest, image, metric report, or index.
Backfill: Execution of a pipeline for historical logical intervals or previously unprocessed partitions.
Batch: In training, the examples processed together before one model update.
Batching: Combining multiple items or requests so one compute operation uses hardware more efficiently.
Blue/green deployment: A release pattern that keeps stable and candidate environments available for a controlled traffic switch.
Block storage: Storage presented to a host as a raw virtual disk, usually with a file system added by the user or operating system.
BM25: A lexical ranking function that scores documents by matching query terms, weighting rare terms more, saturating repeated occurrences, and normalizing for document length.
Calibrated judge review: A model applies a rubric to many context-answer pairs after its verdicts have been compared with cases whose labels human reviewers settled. Disagreement and change still need measurement.
Canary: A candidate release exposed to a small controlled share of live traffic before wider promotion.
Checkpoint: A saved state sufficient to resume a computation under stated compatibility conditions.
Chunk: A source segment used as one embedding and retrieval unit, with links back to its source.
Chunked prefill: Splitting a long prompt into smaller processing chunks so decode work from other requests can run between them.
Classification accuracy: The share whose predicted labels match their reference labels within a nonempty set of evaluated cases.
Classification precision: TP divided by TP plus FP: the share of predicted positive cases that are true positives. It is undefined when no case is predicted positive.
Classification recall: TP divided by TP plus FN: the share of reference-positive cases found by the detector. It is undefined when the reference contains no positive cases.
Cohen’s kappa: Agreement between two labelers adjusted for agreement expected from their label frequencies. It measures agreement, not whether either label is correct, and is undefined when expected agreement is one.
Cold start: Preparation of a new serving replica before it can accept traffic, including required capacity, model loading, initialization, and readiness checks.
Collective operation: A group communication step in which every participating process contributes or receives a defined result, such as a sum.
Concept drift: A change in the relationship between inputs and the target outcome. Its effect on a particular model’s accuracy must be measured.
Consistency: In evaluation, the stability of a verdict across repeated runs of an unchanged case and configuration. A consistency rate can count agreement with that case’s modal label.
Container: A running process created from an image, isolated with operating-system mechanisms while sharing the host kernel.
Container image: A filesystem and configuration template identified by a content digest. A movable tag is not the same fixed identity.
Context parallelism: Distribution of long sequence context across devices to reduce per-device memory or computation.
Continuous batching: Serving that admits waiting sequences and removes finished ones between decode iterations instead of keeping one fixed request batch.
Control plane: The component that stores desired state and reconciles deployments, capacity, configuration, and rollout actions.
Cosine similarity: For two nonzero vectors, the dot product divided by the product of their lengths, comparing direction rather than magnitude.
Critical path: The longest dependency path in a finite DAG. With fixed task durations and finish-before-start dependencies, its duration is a lower bound on completion time. Queues and limited resources can make the run longer.
Data drift: A change in the distribution or schema of model inputs relative to a reference population.
Data parallelism: Replicated computation over distinct input shards with synchronized model updates.
Decode: Autoregressive generation after prefill, usually one next-token step per active sequence.
Deterministic planning allocation: An intentionally assigned time allowance for one stage in a latency design budget. It is a design value, not a measured end-to-end percentile.
Device: One accelerator visible to a training process, such as a GPU.
Direct human semantic review: A person checks an answer’s claims against supplied context, records supporting or contradicting passages, and resolves unclear evidence or policy with reviewers or owners.
Directed Acyclic Graph (DAG): A graph whose directed edges express task dependencies and whose lack of cycles permits an order in which tasks can run.
Distributed Data Parallel (DDP):
PyTorch’s replicated-model training strategy: processes use distinct input shards, synchronize gradients, and apply aligned optimizer updates.Distributed sampler: A policy that assigns dataset indices to data-parallel ranks.
PyTorch’s default padding can repeat indices for uneven datasets. Dropping, worker randomness, transforms, and resume state need explicit policies.2Embedding: A fixed-length numeric vector produced by a specified model and preprocessing pipeline to represent an item for comparison.
Evaluation: Comparing a model’s outputs with reference labels or evidence and recording a result without updating model weights.
ETag: An object-store comparison value for a current object. In the checkpoint protocol, the promotion service supplies the expected pointer ETag in an
If-Matchupdate; pinned reads use object version IDs instead.Evaluator: A rule, model, or human protocol that maps system evidence to a quality judgment.
Experiment run: One traceable execution with recorded parameters, code, inputs, outputs, metrics, and environment.
Expert parallelism: Distribution of mixture-of-experts components and token routing across devices.
Exponential backoff: A retry policy that multiplies the delay between attempts, usually up to a configured limit, to give a temporarily unavailable dependency time to recover. It does not repair a repeatable code or data error.
F1 score: A balance of classification precision and recall, equal to 2TP divided by 2TP + FP + FN. It is undefined when that denominator is zero unless an evaluation policy declares another convention.
False negative (FN): A reference-positive case that a detector predicts as negative.
False positive (FP): A reference-negative case that a detector predicts as positive.
Fine-tuning: Further training of a pretrained model on task-specific examples, either of all weights or of an adapter.
FSDP: Fully Sharded Data Parallel, which shards selected model, gradient, and optimizer state across ranks.
Gang scheduling: Admission of a configured minimum worker group. The complete job must match that configuration, and application readiness is a separate step.
Generation: For a checkpoint save, the fixed set of shards and manifest written for one attempt. A restore reads the approved object versions named by that manifest.
Gradient: The rate at which a training loss changes with a model parameter. Backpropagation computes it through the operations that produced the loss.
Ground truth: A reference label or outcome treated as correct for evaluation, together with its annotation rules and record of how it was established.
Groundedness: The degree to which every claim in an answer is supported by the passages supplied to the model. Support does not establish that the passage itself is true or current.
Hallucination: Generated content that the supplied evidence does not support or that contradicts it.
HNSW: A layered proximity-graph ANN index with tunable construction and search effort.
Hybrid retrieval: Combination of dense, sparse, metadata, and optional reranking stages.
Idempotent task: A pipeline task whose repeated execution with the same declared inputs reaches the same external result without corrupting state or duplicating effects. This property is task idempotency.
Indirect prompt injection: Instructions embedded in retrieved or external content that try to redirect an application. Permission to retrieve the content does not give those instructions authority over application rules.
Inference: Using fixed model weights to turn an input into an output, without a training update.
Inference engine: The runtime that schedules model forward passes, manages memory/KV state, and emits generated tokens.
Infrastructure responsibility: The division of host, networking, runtime, recovery, and other platform work between an application team and a provider.
Inference router: The request-path component that chooses an eligible engine replica using health, queue, model placement, and cache-state evidence.
Inter-Token Latency (ITL): Elapsed time between successive generated tokens during decode, measured over a stated workload.
IOPS: Completed storage input/output operations per second for a stated operation shape.
IVF: An inverted-file ANN method that searches selected coarse vector partitions.
Key-Value (KV) cache: Attention key and value tensors retained from earlier tokens so decode does not recompute the entire prefix.
Large language model (LLM): A model that generates text from supplied input. The running service supplies a question and retrieved passages as that input.
Lineage: The recorded graph connecting outputs to exact inputs, code, configuration, environment, and producer run.
Local rank: One process’s index within its node, commonly used to select the local GPU. It differs from its rank in the full group.
Log: A discrete timestamped event record with structured diagnostic context.
Loss: A number measuring a training prediction against its target for the current batch.
Machine Learning Operations (MLOps): Engineering practice for automating, operating, measuring, and scaling model training, inference, and data pipelines.
Mean reciprocal rank (MRR): The average, over a set of questions, of one divided by the position of the first relevant result. A question with no relevant returned result contributes zero.
Metric: A numeric measurement recorded over time, such as a request count, a derived latency percentile, or the current number of free cache blocks.
Micro-batch: A smaller part of a training batch processed at one time to fit memory or flow through pipeline stages. Several micro-batches can contribute to one model update.
MFU: Model FLOPs Utilization: observed model arithmetic rate divided by relevant hardware peak.
Model registry: A governed store of immutable model versions, lineage, metadata, and promotion state.
NCCL: NVIDIA’s communication library implementing GPU-oriented collective operations.
Normalized discounted cumulative gain (nDCG): Position-discounted relevance of a ranked list divided by the best possible top-\(k\) score from the query’s judged candidates. It is defined here only when that ideal score is positive.
Object storage: API-addressed storage for named objects with semantics distinct from a POSIX file system.
OpenTelemetry (OTel): An observability project with APIs, SDKs, data models, and protocol specifications for traces, metrics, and logs.
OpenTelemetry Collector: A service that receives telemetry, processes or samples it, and exports it to storage or analysis backends.
Optimizer: An update rule that uses gradients, settings, and sometimes saved state to change model parameters during training.
Orchestrator: A system that coordinates when and where work runs and records execution state. Cluster orchestration places workloads on machines. Workflow orchestration coordinates dependencies and attempts across tasks.
Ownership epoch: An increasing number that identifies the coordinator currently allowed to publish a shared pointer or result.
Paged attention: An attention-memory scheme that stores KV state in fixed-size, non-contiguous blocks and maps logical token positions to physical blocks.
Parameter (weight): An adjustable number in a model that training can change.
Pipeline: A reproducible set of tasks and dependencies that transforms versioned inputs into versioned outputs while recording execution state.
Pipeline bubble: Idle time of pipeline-parallel stages while the first micro-batches fill the pipeline and the last ones drain from it.
Pipeline parallelism: Distribution of consecutive model layers across stages that exchange activations and gradients.
Pinned host memory: Host memory kept resident so a device transfer does not need an extra pageable-memory staging step.
Precision@k: The number of relevant results within the top \(k\) divided by a positive cutoff \(k\). Unfilled positions remain in the denominator.
Prediction drift: A change in the distribution of model outputs relative to a reference.
Prefill: Parallel processing of input prompt tokens to create the initial KV state.
Prefill/decode disaggregation: Placing prompt processing and token decoding in different worker pools and transferring KV state between them.
Product quantization: Encoding parts of a vector as compact codes, making stored vectors smaller and computed distances approximate. It can be used with a flat scan or a candidate-pruning index.
Prompt decision tree: An explicit ordered representation of evaluator decisions, overrides, and terminal verdicts.
Quantization: Representation of weights, activations, or cache values with reduced numerical precision.
RadixAttention:
SGLang’s radix-tree index for reusable token prefixes, with unused cached leaves evicted when capacity is needed.Rank: One process identity in a distributed process group, global and local ranks serve different scopes.
Recall@k: For a nonempty reference set, the fraction of relevant reference items found among the top-\(k\) retrieved results.
Reciprocal Rank Fusion (RRF): Combining ranked result lists by each item’s position, without adding raw scores whose scales differ.
Recovery Point Objective (RPO): The maximum acceptable amount of completed work or data lost after failure. For a training run, this is completed training progress.
Recovery Time Objective (RTO): The maximum acceptable time from failure to restored service or useful computation.
Registry: A service that stores versioned container-image layers and supplies them to build or deployment systems.
Release candidate: One immutable serving bundle, retrieval collection, prompt/policy configuration, and evaluation report proposed for production traffic.
Release descriptor: An immutable record of component versions selected to work together. In retrieval it pairs a query encoder with a compatible collection. A full-system descriptor can also identify the serving bundle, prompt and policy configuration, and evaluation evidence. Approval and traffic routing are separate decisions.
Rendezvous: The discovery and coordination mechanism through which distributed ranks form a process group.
Reranker: A more expensive scorer applied to a reduced candidate set to improve final ordering.
Retrieval-augmented generation (RAG): A design in which the application retrieves passages at request time and supplies them to the model as context for the answer.
Server or node: One physical or virtual machine containing CPUs, host memory, storage interfaces, networking, and one or more accelerator devices.
Service Level Objective (SLO): A measurable target for service behavior over a stated population and time window.
Sequence parallelism: Distribution of selected activation operations across the devices of a tensor-parallel group.
Serving bundle: The immutable weights, tokenizer, engine image, configuration, compatibility, and release metadata deployed together.
Shared file storage: Storage that presents one hierarchical file namespace to multiple nodes.
source watermark: A position in an ordered source-change history through which all required updates, permission changes, and deletions have been applied. Matching watermarks show the same change range was applied, not that search results are identical.
Span: A record of one operation, with start and end times and associated details. Linked spans form a trace across components.
Speculative decoding: Generation that proposes tokens with a draft model and verifies them with the target model.
Task specification: The contract for one pipeline task: identity, immutable inputs, outputs, runtime, resources, retry behavior, validation, and side effects.
Tensor parallelism: Distribution of tensor operations within a model layer across devices.
Throughput: Completed useful work per unit time for a stated workload, such as output tokens per second per GPU or requests per second.
Time to First Token (TTFT): Elapsed time from request admission until the first generated token becomes available.
Token: One unit of a tokenizer’s text sequence: a word, part of a word, punctuation, or a special marker, depending on the tokenizer.
Tokenizer: Model-specific software that maps text to token IDs, integer identifiers for vocabulary entries, and maps generated IDs back to text.
Trace: Related spans connected by propagated context that describe one operation across components. The trace records timing and parent-child links, but those links alone do not prove the cause of a failure.
Training: Comparing a model’s prediction with a target and using the resulting objective to update model weights.
True negative (TN): A reference-negative case that a detector predicts as negative.
True positive (TP): A reference-positive case that a detector predicts as positive.
Vector database: A stateful service that stores vector records and metadata and supports similarity-oriented retrieval.
Weight quantization: Storing model weights in fewer bits, reducing weight memory and transfer cost while introducing approximation error that needs task-specific quality checks.
World size: The number of processes participating in a distributed process group.
Zero Redundancy Optimizer (ZeRO): A staged distributed-training method that first shards optimizer state, then gradients, then parameters across data-parallel ranks. ZeRO-3 applies all three stages.
Faiss maintainers. (2025, July 28). Faiss indexes (Faiss project wiki). https://github.com/facebookresearch/faiss/wiki/Faiss-indexes. Faiss documents exhaustive compressed search alongside candidate-pruning indexes. This is the same ANN distinction taught in Section 5.3.↩︎
PyTorch Contributors. (n.d.). torch.utils.data. PyTorch 2.8 documentation. https://docs.pytorch.org/docs/2.8/data.html#torch.utils.data.distributed.DistributedSampler. The documented PyTorch 2.8 default pads uneven datasets. Section 2.3 explains the statistical and repeatability boundary.↩︎