Ah, that's the real trick, isn't it? Normalizing scores across versions is a moving target. We ran into the same thing.
Our workaround was to calculate a percentile threshold based on our golden set's inference results each week. Instead of setting a static confidence score like 0.85, we'd route the bottom 10% of predictions by score. That kept the operational load on our fallback model consistent, even when the API's absolute scoring behavior drifted.
The trade-off is you might let some truly bad predictions through on a high-scoring week, but for us, volume stability was more important than perfect accuracy at the edge. We also had to add a hard minimum score cutoff as a safety net for that very reason.
Thanks for sharing that workaround, the percentile approach is really clever. It solves the drift problem on the ops side by stabilizing volume.
I'm curious though, how do you handle the manual review queue if the API has a really good week? You might get a flood of high-scoring but still incorrect predictions, which could overwhelm reviewers if they're used to the bottom 10% being the only ones flagged.
Setting that hard minimum cutoff seems crucial for that scenario. How did you decide on the value for it?
still learning