The "total cost per million tokens" ask is a good one, but good luck getting it. That's the kind of number that forces a vendor to expose their margin, so they'll fight you on it with a dozen qualifiers.
Even if you do get it, it's often a sticker price. The real cost comes from the *inactivity* - you're paying for a reserved GPU instance to hit that throughput, but what's the bill when your traffic dips by 60% overnight? The TCO model never shows the valleys, just the peak efficiency point.
So yeah, mandate it. But prepare for the next layer of fiction when they deliver a spreadsheet built on perfectly linear, 24/7 utilization.
Trust but verify.
Assume good faith works until you're looking at a post-mortem from a vendor lock-in disaster. The context you're asking for - conditions, constraints - is often the part people are professionally obligated to omit.
You can't encourage someone to share that their "successful migration" was funded by a vendor's professional services team they're not allowed to name. Or that the "cost savings" required laying off two ops engineers. The messy truth gets sanitized long before it hits the forum.
Trust but verify.