Casebook/Case 07
Training & distribution

Scaling Kubernetes to 7,500 nodes

Control-plane, networking, and operational lessons from very large ML clusters.

Reported by the primary sourceFACT LAYER

What we can attribute directly

OpenAI published experience scaling Kubernetes clusters to 7,500 nodes.

The account discusses API load, networking, monitoring, and operational practices.

Read OpenAI — Kubernetes at scale Primary source · last checked 26 Aug 2026
01 · Problem & constraints

The operating envelope

Large node count, bursty jobs, control-plane limits, network fan-out, observability volume, and failure recovery.

Actors

Model teams, platform owners, operators, downstream product systems, and people affected by decisions.

Evidence

Versioned data, configs, traces, artifacts, deployments, and outcomes aligned on one timeline.

Failure cost

Schedulers and control-plane services become a bottleneck before accelerator compute is saturated.

02 · Architecture reconstruction

Trace the system before naming the bug.

  1. 01

    Producers emit versioned data or model artifacts.

  2. 02

    A platform validates, computes, stores, schedules, or routes them.

  3. 03

    Training or inference consumes the exact declared version.

  4. 04

    Telemetry joins the decision to system, data, and model identity.

  5. 05

    Operators compare outcomes, stop conditions, and the last known-good path.

03 · Symptoms & investigation

Follow the evidence boundary by boundary.

Symptoms

Schedulers and control-plane services become a bottleneck before accelerator compute is saturated.

Investigation

Measure API request classes, list/watch behavior, scheduler latency, DNS/network pressure, and telemetry cardinality.

DIAGNOSTIC EXERCISE

A thousand workers start at once. Draw the control-plane request fan-out and identify two ways to smooth it.

Open investigation scaffold
  1. Write the earliest known-bad timestamp.
  2. Compare exact identities on either side of that boundary.
  3. Find the smallest affected slice and a known-good counterexample.
  4. Separate mitigation from root-cause confirmation.
04 · Root cause & fix

Repair the contract, not only the symptom.

ROOT CAUSE

Components designed for smaller steady workloads face burst and fan-out amplification at cluster scale.

FIX

Cache and scope reads, reduce churn, shard or isolate workloads, and design load tests for control-plane behavior.

Rollout

Raise scale in tested steps with explicit capacity headroom and failure drills.

05 · Rejected alternatives

Reason about the tempting shortcuts.

  • Treating the orchestration layer as free overhead.
  • Unlimited high-cardinality cluster telemetry.
06 · Monitoring after the fix

Make recurrence visible early.

01

Scheduling and API latency

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

02

Watch/list volume

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

03

Node and network error budgets

Define owner, slice, normal range, alert persistence, and the exact mitigation the alert should trigger.

REUSABLE PRODUCTION PATTERN

At scale, metadata and coordination become a workload of their own.

Carry this pattern into assignments as a design constraint and into incident reviews as a hypothesis—not as proof about an unpublished system.