The Platform Before the Problems

Decide what you would do from the brief alone, including whether you would change anything at all. Everything below it is available, but the exercise stops working if you open it first.

The brief you were given

A new head of ML platform proposes standardising on a full orchestration and serving platform — pipelines, registry, feature store, serving mesh — with a year of migration for the company's eleven models. The pitch: "We do not have MLOps. We need Kubernetes-native ML." Three of the eleven models have had production incidents this year.

The trap — the fix that moves the metric and is not the fix

Approving the year-long migration and starting with the three incident-prone models "since they need it most". Their migration takes the first two quarters, during which they keep their current pipelines and none of the missing practices; a fourth incident on one of them — a second leaked promotion — is attributed to the old stack. When migrated, the models have a registry and a pipeline runner and still have no promotion gate, because the gate is a test somebody has to write, and the platform team was building infrastructure. The platform is real and the practices still do not exist.

Read this even if you are confident. It is here rather than behind a button because it is the answer most teams actually ship, it passes review, and its cost arrives weeks later when the labels do.