Chapter 1 · Reproduce the workflow

From notebook to reproducible run

Use from notebook to reproducible run to move the engineering foundations production brief toward a defensible release.

65–90 min4 key conceptsReviewed 26 Aug 2026
01 · Production proposition

A churn notebook cannot be reproduced by the teammate expected to ship it.

This lesson isolates from notebook to reproducible run as one decision inside that system. The people affected are product users and operators; the learning data must carry event time, availability time, ownership, and version; and the operating envelope is a declared latency, cost, and operator-capacity budget.

Decision

Choose whether and how to use immutable data snapshots at a declared prediction cutoff.

Metric

Measure decision utility alongside calibration, slice reliability, and system latency—not model score alone.

Failure consequence

A churn notebook cannot be reproduced by the teammate expected to ship it. An unsafe release must degrade to a named baseline or the last known-good version.

02 · Intuition & prerequisites

Build the mental model before the machinery.

The core move is to treat from notebook to reproducible run as a contract between data, a computation, and an action. Reject promotion when any artifact lacks lineage, golden raw-request parity, or a failure-safe recovery path. The implementation becomes easier to debug once you can state which inputs exist, which state is learned, what output means, and what must remain invariant after serialization.

01

immutable data snapshots

Define it in a hand-checkable form and name the prediction-time inputs.

02

dependency locks

Connect it to the production metric and identify what it cannot guarantee.

03

run manifests

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

04

schema contracts

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

Bring forward

Tested Python, basic probability, vectors, and supervised ML.

03 · Formal treatment

Name every symbol. Check every shape.

Content-addressed run identity is the central invariant for this lesson. The formula is useful only when its inputs match the production cutoff and its output maps to an action.

Formal treatment
run_id=SHA256(codedataimageconfigseed)\operatorname{run\_id} = \operatorname{SHA256}(\mathrm{code} \mathbin{\|} \mathrm{data} \mathbin{\|} \mathrm{image} \mathbin{\|} \mathrm{config} \mathbin{\|} \mathrm{seed})

Content-addressed run identity

Symbol, shape or unit contract
SymbolMeaning / shape / unit
codesource revision digest
dataimmutable dataset digest
imageOCI image digest
seeddeclared random seed
Open derivation and numerical substitution

Start from the production quantity being optimized, substitute the observed values with their declared units, then isolate the model-controlled term. Preserve shape annotations at each step so broadcasting or aggregation cannot silently change the result.

  1. Write the named inputs: code, data, image, seed.
  2. Substitute one small, hand-checkable batch before vectorizing.
  3. Calculate an independent reference value and compare within a declared tolerance.
# equation → code contract
inputs = validate_shapes_and_units(batch)
value = compute_c01(inputs)
assert is_finite(value)
04 · Three views of the idea

Calculate it small. Shape it realistically. Break it on purpose.

HAND-CALCULATED TOY

A result you can reproduce on paper

Fit the same line twice from four rows, one locked environment, and one seed; compare manifests and predictions.

  1. Write every input and unit.
  2. Substitute values into the content-addressed run identity equation above.
  3. Compare the result to one simple baseline and explain the direction of the difference.
PRODUCTION-SHAPED

The same reasoning under real constraints

A nightly churn run records source commit, Parquet snapshot, image digest, feature version, metrics, and artifact checksum.

The production record includes the data snapshot, transformation state, artifact identity, cutoff, score, decision, and the version of the policy that consumed it.

FAILURE / COUNTEREXAMPLE

The attractive result you should reject

A mutable data/latest path and unpinned dependencies create a different model from the same command.

Diagnostic: replay the smallest failing slice from immutable inputs, then compare each boundary rather than retuning the model.

05 · Deterministic lab

Change one assumption and make the tradeoff visible.

This lab runs predefined TypeScript only. It never executes learner code. Use the slider, numeric input, reset, live text, or table—the computation is the same.

Simplified simulation

Reproducibility fault injector

Change how many declared inputs remain mutable and observe replay parity.

inputs
Primary72.0%
Secondary3/5 captured
DiagnosisManifest incomplete

Assumption: Each mutable input independently creates a 14-point replay risk.

Open nonvisual data table
ItemComputed stateInterpretation
input 1mutablereplay risk
input 2mutablereplay risk
input 3content-addressedverified
input 4content-addressedverified
input 5content-addressedverified
06 · Production implications

Trace the complete operating path.

  1. 01

    Validate and version immutable data snapshots.

  2. 02

    Compute from notebook to reproducible run from prediction-time-safe inputs.

  3. 03

    Persist model, feature, and configuration identities together.

  4. 04

    Serve or materialize behind explicit a declared latency, cost, and operator-capacity budget.

  5. 05

    Join telemetry to mature outcomes and retain a rollback path.

Observability

Join service health, input quality, prediction distributions, slice behavior, and mature outcomes by exact version.

Cost

Measure storage, preprocessing, compute, queueing, and human review under a representative arrival pattern.

Failure modes

A mutable data/latest path and unpinned dependencies create a different model from the same command. Add a detector, owner, mitigation, and stop condition for this class of failure.

Alternatives

Compare a rule, a simpler statistical baseline, and a different system boundary before adding model complexity.

07 · Check understanding

Explain the contract, not just the vocabulary.

Browser-graded checkpointPass ≥ 80%
01Does reusing a seed make changed data reproducible?
02Why record an image digest?
03What makes a retry safe?
08 · Apply in production

Reproducible baseline package

Design a Docker-only workflow that validates data, trains, evaluates, serializes, performs golden inference, and emits complete lineage.

  • Run manifest and data contract
  • Raw-input golden parity test
  • Failure and retry runbook
  • Model card with residual nondeterminism
Open assignment and rubric
09 · Sources & next depth

Read primary material with a purpose.

10 · Production resolution

Return to the opening failure.

Reject promotion when any artifact lacks lineage, golden raw-request parity, or a failure-safe recovery path.

For this lesson, the release evidence is a hand-checked formal result, deterministic simulation output, a ≥80% checkpoint, the production rubric, and a named fallback. The course resolves when the system can produce a seeded, tested, packaged, versioned training workflow with a model card.