Chapter 1 · Frame and baseline

Choose batch, online, streaming, edge, or hybrid

Use choose batch, online, streaming, edge, or hybrid to move the serving systems production brief toward a defensible release.

65–90 min4 key conceptsReviewed 26 Aug 2026
01 · Production proposition

Launch traffic causes p99 spikes and leaves expensive GPUs idle.

This lesson isolates choose batch, online, streaming, edge, or hybrid as one decision inside that system. The people affected are product users and operators; the learning data must carry event time, availability time, ownership, and version; and the operating envelope is p99 under the product timeout.

Decision

Choose whether and how to use prediction contracts at a declared prediction cutoff.

Metric

Measure decision utility alongside calibration, slice reliability, and system latency—not model score alone.

Failure consequence

Launch traffic causes p99 spikes and leaves expensive GPUs idle. An unsafe release must degrade to a named baseline or the last known-good version.

02 · Intuition & prerequisites

Build the mental model before the machinery.

The core move is to treat choose batch, online, streaming, edge, or hybrid as a contract between data, a computation, and an action. Treat inference as a distributed system with explicit degradation, auditability, rollout, and reversal. The implementation becomes easier to debug once you can state which inputs exist, which state is learned, what output means, and what must remain invariant after serialization.

01

prediction contracts

Define it in a hand-checkable form and name the prediction-time inputs.

02

REST/gRPC

Connect it to the production metric and identify what it cannot guarantee.

03

batching

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

04

timeouts

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

Bring forward

GPU training, Efficient inference, ML platforms

03 · Formal treatment

Name every symbol. Check every shape.

Queue utilization and Little’s law is the central invariant for this lesson. The formula is useful only when its inputs match the production cutoff and its output maps to an action.

Formal treatment
ρ=λcμandN=λW\rho = \frac{\lambda}{c\mu} \quad \text{and} \quad N = \lambda W

Queue utilization and Little’s law

Symbol, shape or unit contract
SymbolMeaning / shape / unit
λarrival rate
cworker count
μservice rate per worker
Wmean time in system
Open derivation and numerical substitution

Start from the production quantity being optimized, substitute the observed values with their declared units, then isolate the model-controlled term. Preserve shape annotations at each step so broadcasting or aggregation cannot silently change the result.

  1. Write the named inputs: λ, c, μ, W.
  2. Substitute one small, hand-checkable batch before vectorizing.
  3. Calculate an independent reference value and compare within a declared tolerance.
# equation → code contract
inputs = validate_shapes_and_units(batch)
value = compute_c18(inputs)
assert is_finite(value)
04 · Three views of the idea

Calculate it small. Shape it realistically. Break it on purpose.

HAND-CALCULATED TOY

A result you can reproduce on paper

At 80 requests/s, two 50 requests/s workers have ρ=.8 yet bursts can still form a long queue.

  1. Write every input and unit.
  2. Substitute values into the queue utilization and little’s law equation above.
  3. Compare the result to one simple baseline and explain the direction of the difference.
PRODUCTION-SHAPED

The same reasoning under real constraints

A typed payment-risk API fetches versioned features, scores XGBoost, logs lineage, and uses a rules fallback.

The production record includes the data snapshot, transformation state, artifact identity, cutoff, score, decision, and the version of the policy that consumed it.

FAILURE / COUNTEREXAMPLE

The attractive result you should reject

Client retries amplify a slow dependency; readiness passes before warmup and newly scaled replicas worsen p99.

Diagnostic: replay the smallest failing slice from immutable inputs, then compare each boundary rather than retuning the model.

05 · Deterministic lab

Change one assumption and make the tradeoff visible.

This lab runs predefined TypeScript only. It never executes learner code. Use the slider, numeric input, reset, live text, or table—the computation is the same.

Estimate

Serving queue lab

Change concurrency and compare utilization, queue pressure, and tail latency.

requests
Primaryρ=0.72
Secondary126 ms p99
DiagnosisHeadroom

Assumption: Four workers, 18 requests/s each, 850 ms timeout, burst factor 1.3.

Open nonvisual data table
ItemComputed stateInterpretation
worker 110 activehealthy
worker 210 activehealthy
worker 310 activehealthy
worker 410 activehealthy
worker queue0 waitinghealthy
06 · Production implications

Trace the complete operating path.

  1. 01

    Validate and version prediction contracts.

  2. 02

    Compute choose batch, online, streaming, edge, or hybrid from prediction-time-safe inputs.

  3. 03

    Persist model, feature, and configuration identities together.

  4. 04

    Serve or materialize behind explicit p99 under the product timeout.

  5. 05

    Join telemetry to mature outcomes and retain a rollback path.

Observability

Join service health, input quality, prediction distributions, slice behavior, and mature outcomes by exact version.

Cost

Measure storage, preprocessing, compute, queueing, and human review under a representative arrival pattern.

Failure modes

Client retries amplify a slow dependency; readiness passes before warmup and newly scaled replicas worsen p99. Add a detector, owner, mitigation, and stop condition for this class of failure.

Alternatives

Compare a rule, a simpler statistical baseline, and a different system boundary before adding model complexity.

07 · Check understanding

Explain the contract, not just the vocabulary.

Browser-graded checkpointPass ≥ 80%
01Why is ρ<1 insufficient for an SLO?
02Shadow versus canary?
03Why not autoscale on CPU alone?
08 · Apply in production

Production inference contract

Specify a typed service with warmup, batching, resilience, load testing, rollout, metrics, and fallback.

  • API/artifact contract
  • Capacity/load report
  • Failure injection
  • Shadow/canary/rollback design
Open assignment and rubric
09 · Sources & next depth

Read primary material with a purpose.

10 · Production resolution

Return to the opening failure.

Treat inference as a distributed system with explicit degradation, auditability, rollout, and reversal.

For this lesson, the release evidence is a hand-checked formal result, deterministic simulation output, a ≥80% checkpoint, the production rubric, and a named fallback. The course resolves when the system can produce batch and online paths with batching, autoscaling, load tests, canaries, and rollback.