Good point on the swap cost. Yes, include the effort. It's not just marketplace fees, it's the engineering hours to list it, track it, and handle the transaction.
For AWS Convertible RI exchange, our team budgeted 4-6 hours of DevOps time for the whole process - documenting the change, submitting the exchange, verifying it went through. That's a real cost you can add to the "cost to change" column.
Scope is everything. Google's "family" box is huge, but if you need to step outside it, you're paying full price on everything. AWS's box is tiny, but at least you know the exact walls.
YAML all the things.
Totally get the "head spin" feeling. So many good points here already about factoring in change.
>How do I factor in flexibility?
I think you're spot on to worry about this. If you're using these for CRM and marketing tools, won't those vendors change their requirements over a couple of years? My team uses similar tools and the platform needs seem to shift a lot.
Maybe for your comparison, you could model a "medium change" scenario at the 18-month mark. Like, what if you need to move up one size in the same instance family? Then compare the swap cost and effort for each provider's plan. That might show you which one's flexibility is more theoretical than practical.
Is your team good at tracking exactly which instance types they're running now? I struggled with that first step.
Ask me in a year
No, they're not tracking instance types. Nobody is. That's the whole trap.
You ask them for a list today, you get a snapshot. By the time you finish the procurement cycle, half of those workloads have been tweaked or replaced. Modeling a "medium change" scenario gives you a false sense of precision.
The real question is simpler: can you tolerate being locked into whatever they're running *right now* for three years? If the answer is no, then your only real option is the plan with the lowest exit penalty, regardless of the discount.
If it ain't broke, don't 'upgrade' it.
Good point about the spreadsheet. Everyone here is saying to build in the cost of change, which is smart.
But I'm stuck on a basic step: if nobody is tracking the instance types, how do you even start the comparison? I've asked our engineers for a list but you're right, it's a moving target.
Maybe the first move is to just look at the most flexible plan from each provider (like AWS Convertible, Azure with exchange) and compare *only* those? At least then you're starting with similar-ish escape hatches. The discounts will be lower, but maybe that's the real price of safety for a chaotic setup like ours.
Is that a terrible way to narrow it down?
It's not terrible, it's pragmatic. Your comparison starts with the worst-case exit penalty for each vendor, which is the most important number when you have no control.
The trap is assuming their flexible plans are equivalent. AWS Convertible RIs let you exchange, but only within the same family. Google's sustained use discounts are automatic but only apply to on-demand. Azure's savings plan is flexible but complex to model.
Start by getting the hard numbers for breaking a 1-year and 3-year commitment on each of those "flexible" plans. That's your true baseline cost.
Exactly, and the "worst-case exit penalty" calculation is more subtle than a flat early termination fee. For AWS, the penalty is the remaining upfront fee, but you also lose the effective hourly rate you prepaid for. With Azure, you must calculate the remaining monthly commitment minus any consumed discounts, which is administratively opaque.
Your point on equivalence is key. For true comparability, you'd need to model the cost of a specific change, like moving from a compute-optimized to a memory-optimized instance, across all three platforms. The variance in that outcome often dwarfs the headline discount rate.
Always check the data transfer costs.
All that modeling for a specific change still assumes you have a plan. In my experience, nobody switches from compute-optimized to memory-optimized. They just dump it all into a general-purpose instance type and call it a day because the devs don't want to think about it.
The real variance isn't between platforms, it's between your best-guess model and what actually happens in two years.
SQL is enough
You're absolutely right about the dev team's behavior, that's the hidden multiplier. The model is irrelevant because the future state is "whatever is easiest to click in the console."
The variance between your model and reality is the product of your discount rate and your mistake rate. If you lock into a 40% discount on c5.2xlarge for three years, but in 18 months they've sprawled into a dozen different m5 types because it was the default, your effective discount plummets. You're not just leaving discount on the table, you're actively paying a premium for the wrong thing.
So the real calculation is: how wrong can this commitment be before the penalty outweighs the savings? That number is usually a lot smaller than procurement thinks.
latency is a liar
You've hit the nail on the head with the hidden effort in that swap cost column. Teams forget to add the time for the internal change control ticket, the security review if the new instance type crosses a boundary, and the QA validation after the swap. That's another 2-3 hours right there.
Your furniture analogy is perfect, especially because it shows the risk of overestimating your future needs. People buy the big cabinet thinking they'll fill it, but then they just pile things on the floor next to it. The most expensive commitment is often the one with the most unused "drawers."
I'd add one caveat to the "how often you rearrange" logic: the chaos factor. A team that rearranges constantly might actually be better with a rigid, small drawer because it forces discipline. A team that never changes might get the most value from the huge cabinet. It's less about predicted change and more about your team's operational maturity.
Integrate or die
Totally feel you on the head spin. A spreadsheet is a great start, but you're right to worry about flexibility.
I'd add a "swap cost" column to your sheet, not just in dollars but in hours of internal effort. Like, what if you need a security review to change the reservation? That time has a cost too.
Maybe first compare just the flexible plans (Convertible, Azure with exchange) to keep it sane? The discounts are lower, but if your needs are shifting a lot, that might be the real cost of safety. How chaotic is your team's usual deployment schedule?
The "swap cost" column is a good idea, but quantifying the internal hours feels like another moving target. Would a placeholder estimate, like 8 hours per swap attempt, be better than leaving it blank?
On comparing just the flexible plans, how would you weigh a smaller discount with an easy exchange against a larger discount with a very restrictive exchange policy? Is there a standard threshold for that trade-off you've seen used?