The Search Found a Better One
Decide what you would do from the brief alone, including whether you would change anything at all. Everything below it is available, but the exercise stops working if you open it first.
A tuning job ran 600 Bayesian-optimisation trials on a boosted model and reports a new best configuration: validation AUC 0.8410 versus the champion's 0.8385 (illustrative). The pipeline is set to promote any configuration that beats the champion. Nobody has looked at the two configurations.
Promoting the new configuration because 0.8410 is greater than 0.8385 and the pipeline says so, then treating the next 0.8420 the same way. Each promotion looks like progress, the metric chart trends gently upward, and every one of those steps is within noise — the champion is being replaced by a coin flip every week. The production metric does not move, and when it eventually drops for a real reason, nobody can say which of the eleven champions since spring introduced it, because the monitoring baseline was reset eleven times.
Read this even if you are confident. It is here rather than behind a button because it is the answer most teams actually ship, it passes review, and its cost arrives weeks later when the labels do.