Spot on about the contract. I've seen teams grumble about a new CloudWatch setup until I showed them the data egress costs hidden in the old vendor's pricing model.
One caveat on the side-by-side: you need to predefine what counts as a "failure" with the team. If you don't, every stutter during the new tool's learning curve gets logged as a strike against it, while the old tool's daily timeouts get written off as "normal."
Ask me about hidden egress costs.
Your point about measuring concrete pain resonates. We applied a similar tactic, but we went one step further and logged *where* the friction occurred. Instead of just counting timeouts, we tagged each event with the user's workflow step using rudimentary screen capture metadata.
The data showed something you'd miss otherwise: 70% of the old tool's timeouts happened during the same three report-building steps. That pattern turned the abstract "it's slow" complaint into a specific engineering ticket to optimize those exact queries in the new system. It made the switch feel like a targeted fix, not a platform gamble.
This is a really practical way to look at it. I like the side-by-side idea, but I'm not sure how to set that up logistically without it becoming a huge manual task. Do you have a script or a simple tool you'd recommend for tracking those timeouts and hallucinations automatically? I can see my team getting frustrated if they have to manually note every single failure.
Also, the contract point is a good hammer. Has anyone actually gotten a vendor to change a clause like that after you pointed it out? Or does it just become the reason you walk away?