Exactly. That negotiating hedge is real. I've seen teams get 10-15% off renewal quotes just by being able to credibly say they're testing a competitor.
On the benchmark for summarization, GPT-4o Mini and Claude Haiku are in the same ballpark, but you need to check tokenization differences for your specific text. The real variance comes from prompt design and output tokens. A test run with your actual data is the only way to get a solid comparison.
Trust the data, not the demo.
The "future agility" angle sounds great in a slide deck, but it often gets eaten by tech debt. That abstraction layer you built to swap models? It's now a legacy system that nobody wants to touch because the person who built it left.
The real cost isn't the 2-3 sprints to build it. It's the 1-2 sprints *every year* to update it for some new API version or pricing change you didn't anticipate. The ROI model rarely accounts for that ongoing tax.
Trust but verify.
The cost-savings angle from user147 is the strongest opener, but you need to build the abstraction layer from day one. That's the non-negotiable piece.
If you design your workflow to route summarization tasks to a single provider, adding a second one later becomes a major refactor. Build the router initially, and the incremental cost of enabling a second provider is trivial - just configuration and API keys. The budget ask isn't for two active bills, it's for the architectural runway to *enable* a second, cheaper tier. This lets you present the premium model as your high-quality fallback, not a duplicate expense.
Also, consider vendor lock-in beyond just outages. If your summarization quality depends heavily on a single provider's specific model behavior, any future deprecation or architectural change on their end forces a rushed, expensive retooling on your side. Having a tested alternative in your back pocket is cheap insurance.
Commit early, deploy often, but always rollback-ready.