Chapter 2 · Build and stress

FSDP and ZeRO-style sharding

Use fsdp and zero-style sharding to move the distributed training production brief toward a defensible release.

65–90 min4 key conceptsReviewed 26 Aug 2026
01 · Production proposition

An eight-GPU job is slower than one GPU and loses progress on failure.

This lesson isolates fsdp and zero-style sharding as one decision inside that system. The people affected are product users and operators; the learning data must carry event time, availability time, ownership, and version; and the operating envelope is eight GPUs with restart under 15 minutes.

Decision

Choose whether and how to use global batch at a declared prediction cutoff.

Metric

Measure decision utility alongside calibration, slice reliability, and system latency—not model score alone.

Failure consequence

An eight-GPU job is slower than one GPU and loses progress on failure. An unsafe release must degrade to a named baseline or the last known-good version.

02 · Intuition & prerequisites

Build the mental model before the machinery.

The core move is to treat fsdp and zero-style sharding as a contract between data, a computation, and an action. Log per-rank data/compute/collective time, preserve complete training state, and keep a one-GPU equivalence test. The implementation becomes easier to debug once you can state which inputs exist, which state is learned, what output means, and what must remain invariant after serialization.

01

global batch

Define it in a hand-checkable form and name the prediction-time inputs.

02

FSDP

Connect it to the production metric and identify what it cannot guarantee.

03

tensor parallelism

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

04

checkpoint resharding

Stress it with a slice, a temporal boundary, and a failure-safe alternative.

Bring forward

Lessons 1, 2 in this course.

03 · Formal treatment

Name every symbol. Check every shape.

Scaling efficiency is the central invariant for this lesson. The formula is useful only when its inputs match the production cutoff and its output maps to an action.

Formal treatment
E=throughputWWthroughput1E = \frac{\mathrm{throughput}_W}{W \cdot \mathrm{throughput}_1}

Scaling efficiency

Symbol, shape or unit contract
SymbolMeaning / shape / unit
Wworker/GPU count
throughput_Wdistributed samples per second
throughput_1single-worker baseline
Open derivation and numerical substitution

Start from the production quantity being optimized, substitute the observed values with their declared units, then isolate the model-controlled term. Preserve shape annotations at each step so broadcasting or aggregation cannot silently change the result.

  1. Write the named inputs: W, throughput_W, throughput_1.
  2. Substitute one small, hand-checkable batch before vectorizing.
  3. Calculate an independent reference value and compare within a declared tolerance.
# equation → code contract
inputs = validate_shapes_and_units(batch)
value = compute_c13(inputs)
assert is_finite(value)
04 · Three views of the idea

Calculate it small. Shape it realistically. Break it on purpose.

HAND-CALCULATED TOY

A result you can reproduce on paper

Two ranks average different local gradients and take identical optimizer steps.

  1. Write every input and unit.
  2. Substitute values into the scaling efficiency equation above.
  3. Compare the result to one simple baseline and explain the direction of the difference.
PRODUCTION-SHAPED

The same reasoning under real constraints

Scale a decoder to eight-GPU DDP, then shard parameters and optimizer state when one GPU cannot fit them.

The production record includes the data snapshot, transformation state, artifact identity, cutoff, score, decision, and the version of the policy that consumed it.

FAILURE / COUNTEREXAMPLE

The attractive result you should reject

One rank skips a batch and misses an all-reduce, leaving every other rank hung.

Diagnostic: replay the smallest failing slice from immutable inputs, then compare each boundary rather than retuning the model.

05 · Deterministic lab

Change one assumption and make the tradeoff visible.

This lab runs predefined TypeScript only. It never executes learner code. Use the slider, numeric input, reset, live text, or table—the computation is the same.

Estimate

All-reduce scaling lab

Add workers and estimate delivered speedup after communication and stragglers.

GPUs
Primary7.0× throughput
Secondary137.8 ms/step
DiagnosisScale validated

Assumption: 120 ms local-compute step; ring communication and logarithmic straggler penalty.

Open nonvisual data table
ItemComputed stateInterpretation
1 workers120.3 ms1.0×
2 workers125.7 ms1.9×
4 workers131.4 ms3.7×
8 workers137.8 ms7.0×
16 workers145.6 ms13.2×
06 · Production implications

Trace the complete operating path.

  1. 01

    Validate and version global batch.

  2. 02

    Compute fsdp and zero-style sharding from prediction-time-safe inputs.

  3. 03

    Persist model, feature, and configuration identities together.

  4. 04

    Serve or materialize behind explicit eight GPUs with restart under 15 minutes.

  5. 05

    Join telemetry to mature outcomes and retain a rollback path.

Observability

Join service health, input quality, prediction distributions, slice behavior, and mature outcomes by exact version.

Cost

Measure storage, preprocessing, compute, queueing, and human review under a representative arrival pattern.

Failure modes

One rank skips a batch and misses an all-reduce, leaving every other rank hung. Add a detector, owner, mitigation, and stop condition for this class of failure.

Alternatives

Compare a rule, a simpler statistical baseline, and a different system boundary before adding model complexity.

07 · Check understanding

Explain the contract, not just the vocabulary.

Browser-graded checkpointPass ≥ 80%
01Why must collectives match in order?
02When prefer DDP to FSDP?
03Why can adding workers change convergence?
08 · Apply in production

Fault-tolerant distributed trainer

Specify DDP correctness, scaling experiments, worker failure, and complete checkpoint recovery.

  • Gradient/data equivalence
  • 1/2/4-GPU efficiency report
  • Kill-and-resume proof
  • Hang/straggler runbook
Open assignment and rubric
09 · Sources & next depth

Read primary material with a purpose.

10 · Production resolution

Return to the opening failure.

Log per-rank data/compute/collective time, preserve complete training state, and keep a one-GPU equivalence test.

For this lesson, the release evidence is a hand-checked formal result, deterministic simulation output, a ≥80% checkpoint, the production rubric, and a named fallback. The course resolves when the system can produce efficient ddp, parallelism choices, elastic sampling, and resumable checkpoints.