Reliability
Reliability in the context of production machine-learning systems.
Lessons
Formal explanations, examples, simulations, and checkpoints.
ML SLIs, SLOs, error budgets, ownership
Use ml slis, slos, error budgets, ownership to move the reliability production brief toward a defensible release.
Reliability, monitoring & experimentationObserve system, data, prediction, and business
Use observe system, data, prediction, and business to move the reliability production brief toward a defensible release.
Reliability, monitoring & experimentationDiagnose, mitigate, and write postmortems
Use diagnose, mitigate, and write postmortems to move the reliability production brief toward a defensible release.
Reliability, monitoring & experimentationRandomized experiments and guardrails
Use randomized experiments and guardrails to move the reliability production brief toward a defensible release.
Reliability, monitoring & experimentationDrift, retraining, rollback, continuous verification
Use drift, retraining, rollback, continuous verification to move the reliability production brief toward a defensible release.
Courses & assignments
Dependency-authoritative learning units.
Casebook
Reported facts and course reconstructions.
D3: automated data-drift detection
Drift detection is triage; impact and diagnosis decide the action.
Cloudflare · Drift & incidentsWhen a generated feature file caused a global outage
Generated data can behave like executable configuration and needs release engineering.
Cloudflare · Drift & incidentsMonitoring ML models for bot detection
Adversarial ML requires joined telemetry and active probes, not passive averages.
Uber · Drift & incidentsRaising the bar on ML deployment safety
ML deployment safety is progressive evidence plus rapid reversibility.