MLOps: From Experiments to Reliable AI Systems

Author

Vitaly Rubinovich

Compilation date
Publisher ID
study-guide-publisher@0.8.14+873a57a

An engineer adjusts a brass rotary dial on a modular tabletop instrument featuring interconnected glowing glass tubes, pressure gauges, a rolling chart recorder, and versioned slide drawers in a workshop laboratory.

Running a model service involves more than producing an answer: the team needs to identify its inputs, handle demand, investigate failures, and check changes before release. This book explains reproducible runs, distributed training, model serving, storage and retrieval, measurement, evaluation, and repeatable workflows. A grounded-answer service connects these topics through concrete operating decisions.

Contents

Front matter

  • Introduction: The Operating Problem
    A model can be impressive in an isolated experiment and still fail as a product. Production introduces traffic, failures, shared resources, changing data, delayed feedback about correctness, cost limits, and deployment risk. When a result changes or a job fails, the team also needs to know which code, data, model, and settings produced it. Together, these constraints turn a one-time result into a system whose behavior and changes the team can explain.

Part I: Reproducible training at cluster scale

An experiment is difficult to repeat or diagnose when its code, data, or runtime is unknown. A larger training run adds other constraints: model state may exceed one GPU, communication can delay updates, and a failed worker can erase useful work. The input data and saved training state must remain available when workers need them.

  • Chapter 1: Making machine learning runs reproducible
    When a model run’s code, data, runtime, or approval record is unclear, another engineer cannot reliably reconstruct what was tested or decide whether that version is ready to deploy. Production adds shared resources, failures, and evidence requirements. Identified artifacts, an explicit responsibility split, and suitable infrastructure turn an ad-hoc run into a repeatable system.
  • Chapter 2: Distributed training across devices and clusters
    How can one training job use many GPUs without turning communication, scheduling, or failure into the real bottleneck? Chapter 1 defined the versioned artifacts and software layers around a run. Single-GPU limits lead to network and process identities, then to recoverable distributed execution and measurement.
  • Chapter 3: Storing training data, checkpoints, and artifacts
    A 512-GPU training run can stall even when every GPU is healthy. If the input path delivers batches late, the GPUs wait. If a save is interrupted halfway, a later restore can load a checkpoint in which some ranks’ shards are missing. Chapter 2 decided what a checkpoint must contain and how often to take one. Whether the run actually keeps its GPUs busy and recovers correctly depends on where its data and checkpoints live and how they are written.

Part II: Reliable model services with retrieval and measurement

A serving system must keep model requests within its time and memory budgets. For the policy-answering service, the application also needs current passages the user may read. Operators then need evidence of whether the model used those passages correctly.

  • Chapter 4: Serving LLMs with predictable performance
    Users of the grounded-answer service expect the first words of an answer within about a second and a steady stream after that. Part I produced a recoverable model checkpoint, but a checkpoint answers no one. Serving replaces the synchronized steps of training with independent requests of different lengths, latency objectives, per-request attention memory, and rolling changes across a fleet of replicas.
  • Chapter 5: Retrieving evidence for grounded answers
    Suppose the remote-work policy changes after the model served in Chapter 4 was trained. Its weights cannot contain that update, but it can still produce a fluent answer. The grounded-answer service finds current passages the user is allowed to read and supplies them with the question. This gives the model evidence for an answer. It does not guarantee that retrieval finds the right passage or that the model uses it faithfully.
  • Chapter 6: Measuring service health and model quality
    The grounded-answer service can return every request in under a second with HTTP status 200 and still tell an employee the wrong remote-work rule. Chapter 4 made the service fast and reversible, and Chapter 5 made it retrieve authorized evidence. Neither shows whether an answer used that evidence faithfully, whether the questions users ask have changed, or which component caused a bad answer.

Part III: Automated workflows and reliable system operation

Independent endpoints, storage systems, and training jobs create coordination work. Automated dependencies, explicit retry behavior, and incident recovery provide one way to manage that work.

  • Chapter 7: Running repeatable ML workflows
    Every week the policy corpus changes. An update can require document validation and new embeddings before the index is rebuilt. Evaluation then informs the release decision. If a step fails halfway or runs twice, an operator needs its recorded inputs and output state to decide how to recover. The same records let another operator continue the work safely.
  • Chapter 8: Operating a grounded-answer system end to end
    Suppose that after a routine weekly update, answers about newly added policies start citing the wrong documents or none at all, while latency and error rates stay normal. The groundedness defined in the Introduction has fallen even though the service is available. The preceding chapters produced versioned artifacts, recoverable distributed jobs, durable storage, serving bundles, a versioned retriever, evaluation evidence, and retryable pipelines. This chapter operates them as one system whose release evidence connects such an alert to a verified recovery.

Appendices

  • Formula reference
    A formula gives a useful estimate only under its assumptions. The table collects relationships for cost, capacity, memory, communication, latency, evaluation, storage, retrieval, and pipeline duration. Each row links to the explanation and worked example that establish its conditions.
  • Questions and concise answers
    A production failure can cross artifacts, compute, storage, serving, retrieval, evidence, and orchestration. The questions follow the chapters in order. Conceptual questions ask which component, evidence, or recovery path explains a symptom or decision. Calculation questions reuse the book’s formulas with new numbers so that the reasoning, not the example values, is tested.
  • Glossary
    MLOps combines vocabulary from infrastructure, distributed systems, model evaluation, storage, retrieval, and orchestration. The same term can connect several components, so each entry identifies the object, behavior, or measurement involved. The entries use consistent meanings across training, serving, evaluation, storage, and pipelines.
  • References
    The sources below support the claims identified by reader footnotes. Entries follow first reader-visible use and list each distinct source once. The footnotes retain the exact support and limits.