I'm currently evaluating Datadog as part of a potential SaaS analytics stack overhaul. Their 'active hour' billing for infrastructure monitoring is confusing me.
The pricing page says we'd be billed per host, per active hour. Does that mean a host that's up for 15 minutes in an hour counts as a full hour? Also, how does this interact with auto-scaling groups? If we have a web server that spins up 5 instances for 20 minutes during a traffic spike, are we billed for 5 full hours?
I'm trying to compare total cost to other vendors, but this model makes it hard to predict. Are there any hidden costs with this approach I should be aware of?
Yeah, their active hour billing can be a bit of a gotcha. From my experience, yes - a host that's up for any amount of time in an hour gets billed for that full hour. So your 15-minute and 20-minute examples would both count as full hours for those hosts.
For auto-scaling groups, it's exactly that kind of spike that makes forecasting tricky. You'll get billed for each of the 5 instances for that hour, even though they only ran for part of it. It's one reason our team ended up setting pretty aggressive scaling cooldowns and looking at custom metrics to trigger scaling, rather than just CPU.
Have you checked if your cloud provider's own basic monitoring (like AWS CloudWatch) would cover your needs for those transient instances? Sometimes using a hybrid approach saves a lot on the Datadog bill.
Clean code is not an option, it's a sanity measure.
It's billed in whole hour increments, yeah. So those short-lived instances will count as full hours.
One thing that helped me forecast was looking at our autoscaling group's hourly instance count metrics in our cloud provider's console first. Gives you a rough idea of the active host-hours you'd be billed for.
Any part of an hour counts, which can add up fast with frequent scaling.
Yes, any partial hour counts as a full billable hour. That's correct for both your 15-minute host and the auto-scaling scenario.
The bigger hidden cost isn't just the spike billing, it's the baseline. If you have hosts that are always on, you're paying for 730 hours per month per host right out of the gate. Compare that to some other tools that bill on averaged daily rates or data volume.
You can model this by pulling your cloud provider's instance uptime metrics for the last month and tallying the total hours where each instance was "active" (any status other than terminated). That sum is your likely Datadog infra bill.