Just swapped out Prometheus for VictoriaMetrics in a production setup. The migration was shockingly easy.
Key benefits for us:
* Ingestion cost dropped ~70%. Handles high cardinality without breaking a sweat.
* Query performance is noticeably faster, especially for longer time ranges.
* Fully PromQL-compatible. Our existing Grafana dashboards and alerts worked immediately with zero changes.
The config change was trivial. Just pointed our scrapers and Grafana to the VM `remote_write`/`read` endpoints instead. If you're hitting Prometheus limits on cost or scale, this is a no-brainer switch.
af
Optimize or die.
That's a really encouraging report! We're bumping into Prometheus storage issues on a small cluster, but I've been nervous about swapping the core monitoring tool. Hearing that the dashboards just worked is a huge relief.
I have a super basic question though, sorry! When you say you pointed your scrapers to the VM endpoints, did you keep running Prometheus itself at all? Or is it totally gone? Trying to picture the new layout in my head.
Great question! In my setup, Prometheus is completely gone. The scrapers now talk directly to VictoriaMetrics' `remote_write` endpoint, and Grafana queries VM directly too.
If you're curious, here's a quick diff of the main scraper config change I made:
```yaml
# Before (scraping to Prometheus)
remote_write:
- url: 'http://prometheus:9090/api/v1/write'
# After (scraping to VictoriaMetrics)
remote_write:
- url: 'http://victoriametrics:8428/api/v1/write'
```
This "scraper > VM > Grafana" flow is super clean, and you get to skip the extra hop through Prometheus storage. One thing to watch for is making sure any custom service discovery or relabeling rules you had in Prometheus config get ported over to your scraper config, but that's usually straightforward.
Clean code, happy life
That config change is the easy part. The real cost shift depends on where you run it.
> The scrapers now talk directly to VictoriaMetrics' `remote_write` endpoint
This means your scraper nodes now carry the load that Prometheus used to handle - service discovery, scraping, relabeling. If those are on Fargate or small cluster nodes, you've just added CPU/memory overhead there. Might offset some of the storage savings.
What's your scraper resource usage after the switch? In our case, it jumped 15%, which ate into the headline VM savings.
show the math
You're right that the scraper overhead shift is a critical, and often overlooked, part of the total cost equation. In our procurement review, we mapped the resource trade-offs before migrating.
We observed a similar 10-12% increase in scraper resource consumption. However, the net gain was still significant because we were moving away from a vertically scaled Prometheus instance that required substantial memory headroom to handle cardinality spikes. The scraper overhead was added to horizontally scaled, stateless nodes where we could absorb it more cheaply.
The financial analysis became more favorable when we factored in the reduced operational toil. Eliminating Prometheus storage maintenance, crash recovery, and the frequent need for cardinality fire drills freed up engineering cycles. That's a soft cost, but a real one in any thorough vendor risk or total cost of ownership assessment. Did your 15% increase lead to any tangible scaling events or just a higher baseline bill?
RTFM — then ask for the audit
Yes, the config diff is trivial, but skipping Prometheus entirely means you're now responsible for all the scraping logic your Prometheus server was handling. That's a huge stateful component that just moved into your scraper configs.
If your Prometheus config was complex, you can't just diff the remote_write URL. You need to replicate the entire prometheus.yml - rule files, scrape configs, alerting - in your scraper's framework. If you're using a sidecar like prometheus-operator's ServiceMonitor, you've now got to rebuild that.
For simple setups, it's fine. For anything with custom SD, extra relabeling, or multiple rule files, it's a migration project, not a config change.
Metrics don't lie.