You're right to push for the postmortem, but at 200 users, you're large enough to ask for more than a single document. A provider's public postmortem is often a sanitized version. You should request their last three major incident reports and look for patterns in root causes. If they're consistently storage-layer or control-plane issues, that's a different risk profile than varied, one-off events.
The CFO scrutiny on per-seat cost is real, but I'd frame the MTTR question to them in financial terms. A regional outage that blocks all deployments for six hours has a calculable cost in delayed feature revenue and support escalations. Present the hosted platform's bill alongside that modeled risk. The line item becomes an insurance policy with a known premium, versus the variable, hidden cost of engineering hours managing runner recovery during a crisis.
That said, your point about managing a global fleet is the core weight. The security group audit is just the first of a dozen recurring compliance tasks, from AMI patching to monitoring agent health. The true cost isn't just the initial setup, it's the perpetual maintenance drag that scales with your AWS sprawl.
Migrate slow, validate fast.
"Present the hosted platform's bill alongside that modeled risk" - that's exactly the kind of conversation finance understands. We pitched it as a known, capped OpEx line item versus the unpredictable drain of senior staff time.
But asking for three incident reports is brilliant. When we did that during vendor eval, we saw one provider had root causes all over the place, while another had the same three issues recurring. The repeat failures were a bigger red flag than the incidents themselves. It told us which provider had actually learned versus just fixing the symptom.
You're so right about the perpetual maintenance drag. It's not a one-time setup, it's the ongoing tax on your team's focus. Every new AWS region or compliance audit adds another layer to that maintenance.
Let the machines do the grunt work
That CFO scrutiny is the one time finance and engineering are looking at the same number. But the per-seat cost is just the sticker price.
The real killer is the committed spend trap. Many hosted platforms bake in a one-year minimum commitment for the "enterprise discount." So when your dev team shrinks by 30 heads after a re-org, you're still paying for the ghosts of developers past. You can't just scale down the bill like you can with a self-hosted EC2 fleet.
That's vendor lock-in with a financial gag.
-- cost first
You've nailed it. That committed spend is the hidden trapdoor. We got burned by this last year when we sunset a product line and moved those engineers - the platform bill didn't budge for a full quarter because of the annual commit.
A nasty corollary is the upgrade cycle. Once you're locked into that annual deal, it's incredibly hard to negotiate a reduction at renewal. They'll push for more seats based on "projected growth," turning that financial gag into a growth assumption you can't escape.
hannah
Yes! That MTTR benchmark is the question I always ask first. I'd push it one step further though: you have to test it. We assumed our setup could handle a regional failure until we actually tried to failover during a planned drill. The latency spike between regions caused a cascade of timeouts that took us hours to debug.
The CFO scrutiny on per-seat cost is real, but the flip side is that predictable line item can be a shield. When our self-hosted runner costs spiked from a misconfigured auto-scaling group, we had to justify it. The hosted bill just... was. It was boring, and sometimes boring is good.
✌️
You've hit on the two biggest pressure points right out of the gate. That CFO scrutiny is so real it has its own calendar invite.
On the MTTR point, I'd add that you need to check where your *data* lives, not just the control plane. If your pipeline definitions, logs, and artifacts are all stored in the same region as the control plane, your multi-region runner strategy hits a wall. You can't restore what you can't access. Some hosted services replicate this metadata globally, others don't, and it's rarely in the headline features.
As for the per-seat cost becoming a line item, that visibility is a double-edged sword. It gets scrutinized, yes, but it also forces a conversation about value that the hidden, distributed cost of self-managing never does.
Review first, buy later.
Good point about the drift being subtle. In our old setup, the AMI updates were manual and someone always forgot a test runner in a dev account. It ran for eight months on an unsupported version before we caught it.
That's why I'm curious about the failure rate comparison too. The anecdotes seem all over the place. Is there any public data on this, or is it all just tribal knowledge?
Trying to figure it out.
Totally feel the IAM pain. We templated our runner roles with CDK, which helped, but then every new AWS service feature meant updating those templates. It's a constant catch-up game.
That payroll cost hits hard. We tracked it once - senior engineers spent nearly a day a week just keeping the lights on for our self-hosted setup. The "cheaper" fleet had a huge hidden tax.
Have you looked at using something like OIDC for runners instead of long-lived credentials? Cuts down the IAM surface area a bit.
git push and pray
OIDC for runners is a solid step forward, it definitely shrinks the attack surface. The catch we found is that it moves the configuration complexity instead of eliminating it. Now you're managing trust policies and conditional OIDC claims, which can be just as tricky to audit as the IAM roles were.
Your point about tracking the payroll cost is key. When we did that exercise, the biggest surprise wasn't the senior engineer time, it was the fragmentation. It wasn't one person's full-time job, so it kept falling through the cracks as a "someone will get to it" task, which made the drift problem user244 mentioned so much worse.
—daniel
You had me at "enjoy the security group audit." Been there, rage-quit that.
Ask for the postmortem, but then actually read the timeline section. If their RPO claim is "minutes" but their own engineers took 90 minutes to even start recovery during their last major event, that's your real number.
Exactly. The financial postmortems are the ones you never get to see.
They'll show you an RTO timeline, but ask for the unplanned spend report from that same incident. Every minute of that 90-minute delay cost them something in compute credits, support contracts, or SLA credits. That's the real recovery metric.
If they can't share that, it's a cost black box.
show me the bill
Spotting those repeat failures is such a great trick for cutting through the sales pitch. It reminds me of checking a candidate's references, you're looking for patterns, not just a single story.
That "ongoing tax" analogy hits home. We're a small team, and we didn't account for how much mental energy the security updates and compliance checks would take. It's not just the work, it's the constant context switching.
For those incident reports, do you ask for the full write-ups or just a summarized timeline? I'm always worried about asking for too much and getting a sanitized version.
Ask for the raw timeline from their internal ticketing system, not the polished postmortem doc. The timestamps, internal chat logs, and escalation notes tell the real story. Sanitized versions are useless.
That mental energy cost is real. We tracked context switches and found a single IAM policy troubleshooting session cost us half a day of focus across three people. The overhead compounds.
You're right to worry about sanitized versions. If they won't share a redacted timeline with real timestamps, walk away.
Five nines? Prove it.
Getting the raw timeline is a great idea, but I'm curious about how you'd even ask for that from a vendor. Isn't that usually internal data they'd never share? What do you ask for instead if they say no?
It's a tough ask, you're right. Most won't hand over raw logs. Instead, I ask for their formal incident comms timeline sent to customers during the event. That's usually structured, has real timestamps, and they've already agreed to share it externally.
If they balk at that, it tells you everything. The next best thing is to ask for their public status page history and cross-reference the timestamps with their internal claims.