Casebook/Case 16
Drift & incidents

D3: automated data-drift detection

Detecting changes across many features without turning every statistical difference into an incident.

Reported by the primary sourceFACT LAYER

What we can attribute directly

Uber published D3 as an automated system to detect data drift.

The system addresses detection across datasets and feature distributions.

Read Uber Engineering — D3 Primary source · last checked 26 Aug 2026
01 · Problem & constraints

The operating envelope

Many models and features, seasonality, multiple comparisons, limited labels, alert fatigue, and ownership.

Actors

Model teams, platform owners, operators, downstream product systems, and people affected by decisions.

Evidence

Versioned data, configs, traces, artifacts, deployments, and outcomes aligned on one timeline.

Failure cost

Silent input changes degrade models, while naive detectors generate too many unactionable alerts.

02 · Architecture reconstruction

Trace the system before naming the bug.

  1. 01

    Producers emit versioned data or model artifacts.

  2. 02

    A platform validates, computes, stores, schedules, or routes them.

  3. 03

    Training or inference consumes the exact declared version.

  4. 04

    Telemetry joins the decision to system, data, and model identity.

  5. 05

    Operators compare outcomes, stop conditions, and the last known-good path.

03 · Symptoms & investigation

Follow the evidence boundary by boundary.

Symptoms

Silent input changes degrade models, while naive detectors generate too many unactionable alerts.

Investigation

Compare reference selection, sample volume, effect size, seasonality, slice behavior, and upstream changes.

DIAGNOSTIC EXERCISE

Feature PSI rises after a holiday. What evidence separates expected seasonality from a broken pipeline?

Open investigation scaffold
  1. Write the earliest known-bad timestamp.
  2. Compare exact identities on either side of that boundary.
  3. Find the smallest affected slice and a known-good counterexample.
  4. Separate mitigation from root-cause confirmation.
04 · Root cause & fix

Repair the contract, not only the symptom.

ROOT CAUSE

A statistically significant shift is not automatically a harmful or actionable production change.

FIX

Combine robust detectors with magnitude, persistence, model sensitivity, lineage, and routing to owners.

Rollout

Backtest on known changes, begin as a dashboard, then alert only on high-precision patterns.

05 · Rejected alternatives

Reason about the tempting shortcuts.

  • Retraining on every drift signal.
  • One universal threshold for every feature.
06 · Monitoring after the fix

Make recurrence visible early.

01

Drift magnitude and persistence

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

02

Alert precision and owner response

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

03

Outcome quality when labels mature

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

REUSABLE PRODUCTION PATTERN

Drift detection is triage; impact and diagnosis decide the action.

Carry this pattern into assignments as a design constraint and into incident reviews as a hypothesis—not as proof about an unpublished system.