User IDs gave you a map of the crime scene after the robbery. The problem was letting a buggy script generate thousands of billable traces in the first place.
Usage-based pricing makes every mistake a financial crisis. You didn't need better attribution, you needed a system that isn't built to profit from your own waste.
If it ain't broke, don't 'upgrade' it.
What you've uncovered isn't an argument for granular attribution, it's the strongest case against consumption pricing I've seen. The vendor's model intentionally made your minor bug maximally expensive. You were charged for every single redundant trace, and their pricing structure ensured there was no natural brake on the waste.
Now you're stuck implementing throttles, reviewing instrumentation, and adding process gates to prevent your own tools from bankrupting you. That's the real cost of ownership they don't put on the sales page. You've become an enforcer for their profit margin.
The next question is whether you're willing to pay a premium for the privilege of policing your internal workflows, or if this is the catalyst to explore systems where a bug is just a bug, not an invoice.
Skeptic by default
The immediate win here is you can now show QA their own bill. That's the one thing that actually changes behavior.
But you've uncovered a vendor pricing incentive that's working against you. Their "per trace" model made your flawed loop infinitely scalable from a cost perspective. That's not an accident.
You've traded a blended monthly line item for the privilege of seeing, in real time, how their pricing structure turns your bugs into their revenue. Are you planning to renegotiate based on this new data, or is this just for internal austerity?
Show me the bill
So you're celebrating the visibility that showed you the scale of the fire. Fine. But the alert you're recommending just monitors the blaze. It doesn't stop the arsonist.
My caveat with the "alert on spikes" approach is it creates an ops burden to police the vendor's pricing model. You've now assigned someone to watch the meter spin. That's a hidden cost shift from their infrastructure to your team's attention.
The real question isn't about catching the next spike faster. It's why your vendor's pricing has no circuit breaker built in.
Trust but verify.