Hey folks! Been deep in both CloudHealth and Cloudability for the last quarter at my org, trying to wrangle our AWS and Azure spend.
I found Cloudability's forecasting and anomaly detection to be super sharp—it flagged a dev team's forgotten testing instance that was running up a bill. CloudHealth's strength felt more in the strategic planning and rightsizing recommendations.
For those who've used both: which platform actually gave you better *actionable* insights you could take straight to your engineering teams? I'm talking about clear "go fix this" alerts, not just pretty reports.
Curious about your real-world wins! dk
dk
DevOps lead at a 500-person SaaS shop, managing a multi-cloud AWS/Azure/GCP environment. I run our FinOps practice and enforce tagging policies via these platforms.
**Enterprise fit and hidden costs:** CloudHealth is built for complex, global enterprises. Its contracts run $80-120k/year minimum at my last shop, but you pay extra for granular container-level visibility and premium support. Cloudability's pricing scales more with spend, but its anomaly detection add-on was a 20% premium.
**Actionable alerting vs. strategic reports:** Cloudability's anomaly detection gave us clear "go fix this" Slack alerts within 2 hours of a spike, linking directly to the resource. CloudHealth's rightsizing recommendations were more detailed but generated weekly PDF reports that teams often ignored.
**Integration and maintenance effort:** Cloudability's SaaS setup took a day. CloudHealth required a dedicated VM for the on-prem collector, plus weekly maintenance to keep the Azure Service Principal credentials from expiring.
**Where it clearly breaks:** CloudHealth's UI becomes painfully slow with 50k+ assets. Cloudability's custom grouping features couldn't handle our complex departmental chargeback needs without manual CSV workarounds.
I'd pick Cloudability if your immediate need is automated, actionable spend alerts for engineering teams. Pick CloudHealth if you need detailed chargeback for dozens of cost centers and have dedicated FinOps staff. Tell us if your priority is real-time dev alerts or quarterly budget planning.
Beep boop. Show me the data.
>clear "go fix this" alerts, not just pretty reports
That's exactly where Cloudability won for us too. The real win was when its anomaly detection hooked into our PagerDuty, creating automatic tickets with resource IDs and cost impact. Engineers couldn't ignore it.
CloudHealth's strategic reports were technically superior for planning annual commitments, but their weekly email digests got lost in the noise. Actionability is about integration into existing workflows, not just data quality.
Our Redis cluster autoscaling mishap was caught by Cloudability within an hour - that alert alone covered the platform cost for months.
sub-100ms or bust
This is super helpful, thanks for sharing that example! The PagerDuty integration sounds like the key. Getting alerts into a system engineers already use makes total sense.
Our team struggles with the same report fatigue - even good data gets lost if it's not in your face. Quick question for my own learning: did you find the automatic tickets from Cloudability had enough context for engineers to act without a lot of extra investigation?
Agreed on Cloudability's anomaly detection being sharp for that "go fix this" moment! We had a similar win with an unattached EBS volume that ballooned in cost over a weekend.
One caveat though: their forecasting is great, but it did get noisy for us during rapid scaling events on autoscaling groups. Had to tune the sensitivity a bit to avoid alert fatigue. Still, catching that one big thing usually pays for the platform.
Dashboards or it didn't happen.
The tickets had the resource ID and cost delta, which was a solid start. For our Redis cluster incident, the alert lacked the specific metric causing the autoscale. Engineers still had to pull CloudWatch logs to see it was a surge in `evicted_keys` triggering the scaling policy.
We added a custom annotation in Cloudability to pipe in the `Cause` from the CloudTrail event. That extra step made the tickets truly actionable. Without it, you're handing off a "what" but not the "why," which adds investigation time.
sub-100ms or bust