Latency
Latency in the context of production machine-learning systems.
Lessons
Formal explanations, examples, simulations, and checkpoints.
Choose batch, online, streaming, edge, or hybrid
Use choose batch, online, streaming, edge, or hybrid to move the serving systems production brief toward a defensible release.
Serving & inference systemsPackage model, preprocessing, schema, and runtime
Use package model, preprocessing, schema, and runtime to move the serving systems production brief toward a defensible release.
Serving & inference systemsFeatures, batching, caching, and backpressure
Use features, batching, caching, and backpressure to move the serving systems production brief toward a defensible release.
Serving & inference systemsCapacity, load testing, autoscaling, and tails
Use capacity, load testing, autoscaling, and tails to move the serving systems production brief toward a defensible release.
Serving & inference systemsShadow, canary, rollback, fallback, observability
Use shadow, canary, rollback, fallback, observability to move the serving systems production brief toward a defensible release.
Courses & assignments
Dependency-authoritative learning units.
Casebook
Reported facts and course reconstructions.
Scaling Kubernetes to 7,500 nodes
At scale, metadata and coordination become a workload of their own.
Shopify · Serving & costA platform for real-time predictions
Online inference is a user-facing distributed system with a model inside.
Meta · Serving & costTail utilization in ads inference
Distributed inference capacity is constrained by the busiest relevant slice, not the average box.