Casebook/Case 20
Drift & incidents

Off-policy correction for recommender feedback

Learning and evaluating from feedback collected by an older ranking policy.

Reported by the primary sourceFACT LAYER

What we can attribute directly

Google Research published top-K off-policy correction for a REINFORCE recommender system.

The work addresses policy mismatch in recommendation learning.

Read Google Research — Top-K off-policy correction Primary source · last checked 26 Aug 2026
01 · Problem & constraints

The operating envelope

Biased exposure, top-K slates, changing policies, high-variance estimators, and online risk.

Actors

Model teams, platform owners, operators, downstream product systems, and people affected by decisions.

Evidence

Versioned data, configs, traces, artifacts, deployments, and outcomes aligned on one timeline.

Failure cost

Offline training amplifies what the logging policy already showed and underestimates unseen alternatives.

02 · Architecture reconstruction

Trace the system before naming the bug.

  1. 01

    Producers emit versioned data or model artifacts.

  2. 02

    A platform validates, computes, stores, schedules, or routes them.

  3. 03

    Training or inference consumes the exact declared version.

  4. 04

    Telemetry joins the decision to system, data, and model identity.

  5. 05

    Operators compare outcomes, stop conditions, and the last known-good path.

03 · Symptoms & investigation

Follow the evidence boundary by boundary.

Symptoms

Offline training amplifies what the logging policy already showed and underestimates unseen alternatives.

Investigation

Reconstruct exposure propensities, policy versions, support overlap, weight distribution, and variance by slice.

DIAGNOSTIC EXERCISE

A new item has zero logging-policy exposure. Can inverse propensity scoring evaluate it? Explain the support condition.

Open investigation scaffold
  1. Write the earliest known-bad timestamp.
  2. Compare exact identities on either side of that boundary.
  3. Find the smallest affected slice and a known-good counterexample.
  4. Separate mitigation from root-cause confirmation.
04 · Root cause & fix

Repair the contract, not only the symptom.

ROOT CAUSE

Observed clicks are conditional on the previous policy’s exposure, not unbiased relevance labels.

FIX

Use logged propensities and off-policy correction with clipping and overlap diagnostics.

Rollout

Validate estimators on controlled policy changes before using them for candidate promotion.

05 · Rejected alternatives

Reason about the tempting shortcuts.

  • Treating non-exposure as a negative label.
  • Unbounded inverse-propensity weights.
06 · Monitoring after the fix

Make recurrence visible early.

01

Policy overlap and effective sample size

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

02

Weight tails and estimator variance

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

03

Online guardrails by slate position

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

REUSABLE PRODUCTION PATTERN

Feedback data identifies policy-conditioned behavior; correction needs overlap and controlled variance.

Carry this pattern into assignments as a design constraint and into incident reviews as a hypothesis—not as proof about an unpublished system.