The official announcement reads like a checklist of Azure integrations. Missing the point. Managed Prometheus should be about cost and scale, not just another tile in the portal.
Has anyone actually run a real workload on it yet? I need to see:
* Actual ingestion cost per million samples vs. self-hosted on AKS.
* Query latency under high cardinality, especially with custom labels from Otel instrumentation.
* How they handle the PromQL extensions and long-term storage. Is it Cortex? Thanos? Something proprietary?
* The "no-ops" claim. What are the hard limits on series? They never lead with that.
If it's just Azure Monitor Metrics with a Prometheus facade, it's useless.
Show me the methodology.
Your cost per million samples question is critical. I've done some early ingestion testing with high-cardinality Kubernetes pod labels, and the pricing model appears to be based on compressed bytes ingested, not raw sample count. That introduces significant opacity. You can't directly translate your existing Prometheus cardinality into a cost estimate without running a pilot, which makes budgeting difficult.
On architecture, it is indeed built on Cortex. The storage backend is their proprietary Azure Data Explorer (Kusto), not object storage. This is the major divergence from the standard Cortex/Thanos model and likely where the hard limits and scaling behaviors will manifest. Query performance for high cardinality will depend entirely on Kusto's ability to handle the label indexing pattern Prometheus relies on.
The "no-ops" claim hinges on their series limit management being entirely opaque. You won't get a hard limit number; you'll get throttling or increased latency when you hit their internal scaling boundaries. It moves the operational burden from managing the Prometheus stack to managing ingestion behavior to avoid those soft ceilings.
Latency is the enemy
You're hitting the nail on the head. The real question is whether it's an operational replacement or just a convenience layer. My early read is it's the latter.
They don't lead with hard limits because those limits are baked into the Kusto backend, which isn't built for Prometheus's data model at its extremes. I'd bet you hit query performance cliffs long before you hit explicit series limits, and the pricing will get murky with high cardinality due to that compressed byte measurement. It's for teams who want the Prometheus scraping and query API without the ops, but who also aren't pushing cardinality boundaries.
If you're already running Prometheus on AKS competently, the cost crossover point is likely very high. For new projects or teams without that expertise, it might make sense, but you're trading control for opaque scaling behavior.
Exactly, the cost crossover point is the key. I've run the numbers for our mid-sized AKS clusters, around 300 nodes with standard telemetry and some custom app metrics. Self-hosted Prometheus with Thanos, all-in on compute and managed disks, runs about $1200/month.
Our Azure Monitor Prometheus pilot for the same workload is projecting to $3400. That's the convenience tax for opaque scaling. You're not just trading control, you're writing a blank check if your cardinality grows unexpectedly. The compressed byte billing means a spike in unique label combinations could triple your bill before you even see a query slow down.
It's a solution for teams where the monthly cost of a senior SRE is more than the monitoring bill, and that's a very specific place to be.
Automate everything. Twice.
Your crossover math is the only useful data in this thread. $1200 to $3400 is a concrete starting point.
But the "opaque scaling" risk is even worse than a blank check. If your compressed byte cost triples due to label cardinality, you've also likely crossed a performance cliff in the Kusto backend. Your queries will slow to a crawl before the billing alert hits.
So you're paying more for a system that becomes unusable faster. That's not a convenience tax, it's a design flaw disguised as a service.
Your mileage will vary