Chapter 3 · Resolve and operate

Drift, retraining, rollback, continuous verification

Use drift, retraining, rollback, continuous verification to move the reliability production brief toward a defensible release.

65–90 min2 key conceptsReviewed 26 Aug 2026
01 · Production proposition

Model quality decays while every infrastructure dashboard stays green.

This lesson isolates drift, retraining, rollback, continuous verification as one decision inside that system. The people affected are product users and operators; the learning data must carry event time, availability time, ownership, and version; and the operating envelope is release error budget and label maturity.

Decision

Choose whether and how to use guardrails at a declared prediction cutoff.

Metric

Measure decision utility alongside calibration, slice reliability, and system latency—not model score alone.

Failure consequence

Model quality decays while every infrastructure dashboard stays green. An unsafe release must degrade to a named baseline or the last known-good version.

02 · Intuition & prerequisites

Build the mental model before the machinery.

The core move is to treat drift, retraining, rollback, continuous verification as a contract between data, a computation, and an action. Every release needs causal evidence, multi-layer observability, tested rollback, and a retraining policy that does not automate faults. The implementation becomes easier to debug once you can state which inputs exist, which state is learned, what output means, and what must remain invariant after serialization.

01

guardrails

Define it in a hand-checkable form and name the prediction-time inputs.

02

champion/challenger

Connect it to the production metric and identify what it cannot guarantee.

Bring forward

Lessons 1, 2, 3, 4 in this course.

03 · Formal treatment

Name every symbol. Check every shape.

Difference-in-means experiment estimate is the central invariant for this lesson. The formula is useful only when its inputs match the production cutoff and its output maps to an action.

Formal treatment
τ^=YˉTYˉC,SE=sT2nT+sC2nC\hat{\tau} = \bar{Y}_T - \bar{Y}_C,\quad \mathrm{SE} = \sqrt{\frac{s_T^2}{n_T} + \frac{s_C^2}{n_C}}

Difference-in-means experiment estimate

Symbol, shape or unit contract
SymbolMeaning / shape / unit
Ȳ_T, Ȳ_Ctreatment/control means
sample variance
nindependent randomized units
Open derivation and numerical substitution

Start from the production quantity being optimized, substitute the observed values with their declared units, then isolate the model-controlled term. Preserve shape annotations at each step so broadcasting or aggregation cannot silently change the result.

  1. Write the named inputs: Ȳ_T, Ȳ_C, s², n.
  2. Substitute one small, hand-checkable batch before vectorizing.
  3. Calculate an independent reference value and compare within a declared tolerance.
# equation → code contract
inputs = validate_shapes_and_units(batch)
value = compute_c19(inputs)
assert is_finite(value)
04 · Three views of the idea

Calculate it small. Shape it realistically. Break it on purpose.

HAND-CALCULATED TOY

A result you can reproduce on paper

Simulate a two-arm conversion experiment and change variance and sample size.

  1. Write every input and unit.
  2. Substitute values into the difference-in-means experiment estimate equation above.
  3. Compare the result to one simple baseline and explain the direction of the difference.
PRODUCTION-SHAPED

The same reasoning under real constraints

Pre-register a user-level ranking test with conversion primary and latency, complaints, and diversity guardrails.

The production record includes the data snapshot, transformation state, artifact identity, cutoff, score, decision, and the version of the policy that consumed it.

FAILURE / COUNTEREXAMPLE

The attractive result you should reject

Daily peeking and session-level analysis of a user-randomized experiment produce false certainty while a regional regression hides.

Diagnostic: replay the smallest failing slice from immutable inputs, then compare each boundary rather than retuning the model.

05 · Deterministic lab

Change one assumption and make the tradeoff visible.

This lab runs predefined TypeScript only. It never executes learner code. Use the slider, numeric input, reset, live text, or table—the computation is the same.

Estimate

Drift triage lab

Increase shift magnitude and compare statistical signal with estimated business impact.

σ
Primary0.11 utility Δ
Secondary73.3%
DiagnosisObserve persistence

Assumption: Ten-thousand observations; model sensitivity fixed at 0.18 utility points per σ.

Open nonvisual data table
ItemComputed stateInterpretation
global0.60σ10.8%
new users0.67σ12.7%
mobile0.74σ14.7%
region A0.82σ16.6%
region B0.89σ18.6%
06 · Production implications

Trace the complete operating path.

  1. 01

    Validate and version guardrails.

  2. 02

    Compute drift, retraining, rollback, continuous verification from prediction-time-safe inputs.

  3. 03

    Persist model, feature, and configuration identities together.

  4. 04

    Serve or materialize behind explicit release error budget and label maturity.

  5. 05

    Join telemetry to mature outcomes and retain a rollback path.

Observability

Join service health, input quality, prediction distributions, slice behavior, and mature outcomes by exact version.

Cost

Measure storage, preprocessing, compute, queueing, and human review under a representative arrival pattern.

Failure modes

Daily peeking and session-level analysis of a user-randomized experiment produce false certainty while a regional regression hides. Add a detector, owner, mitigation, and stop condition for this class of failure.

Alternatives

Compare a rule, a simpler statistical baseline, and a different system boundary before adding model complexity.

07 · Check understanding

Explain the contract, not just the vocabulary.

Browser-graded checkpointPass ≥ 80%
01Why not retrain automatically on any drift?
02What should an ML SLO cover?
03Why match analysis to randomization unit?
08 · Apply in production

Launch analysis and incident response

Analyze a synthetic launch, verify assignment, estimate effects, inspect slices, and choose continue, pause, or rollback.

  • Experiment validity report
  • Layered reliability diagnosis
  • SLO and runbook
  • Postmortem/retraining policy
Open assignment and rubric
09 · Sources & next depth

Read primary material with a purpose.

10 · Production resolution

Return to the opening failure.

Every release needs causal evidence, multi-layer observability, tested rollback, and a retraining policy that does not automate faults.

For this lesson, the release evidence is a hand-checked formal result, deterministic simulation output, a ≥80% checkpoint, the production rubric, and a named fallback. The course resolves when the system can produce joined data/model/service telemetry, drift policy, experiments, alerts, and runbooks.

Course production assignmentLaunch analysis and incident response