The Dashboard Says 100%
Decide what you would do from the brief alone, including whether you would change anything at all. Everything below it is available, but the exercise stops working if you open it first.
A newly launched credit model's dashboard has shown "accuracy 100%" since launch. The team celebrated and moved on. Labels for default arrive 30 days after the first payment date, which is itself 30 days after approval. Sixty-five days after launch, the first defaults are arriving and the number is falling by the hour. Leadership asks whether the model is failing.
Retraining immediately on the maturing labels "to fix the falling accuracy". The retrain uses a training set in which the only labelled recent applicants are the first cohort's defaults — a small, early-maturing, unrepresentative sample — and the model shifts toward declining applicants who look like them. The dashboard number, now computed on a tiny matured set, jumps around and eventually looks better because the retrained model declines the population that produced the early labels. Nothing has been learned about whether the launched model was working, and the retraining system has now been taught to react to a dashboard artefact.
Read this even if you are confident. It is here rather than behind a button because it is the answer most teams actually ship, it passes review, and its cost arrives weeks later when the labels do.