Alright, so I've officially rotated out of my last observability vendor (we'll call them "Vendor X" for now, NDA and all that) and decided to give Grafana Cloud's free tier a proper shakedown. My usual M.O.: take something they claim is for "small production loads" and actually put a modest, but real, production workload on it. See what bends, what breaks, and where the invisible fence really is.
I'm running a suite of about a dozen microservices, doing maybe 50k metrics per minute, a few hundred spans per minute, and forwarding logs from a couple of critical app containers. The pitch of "forever free" up to 50GB of logs, 50GB of traces, and 10k metrics series is alluring, especially after dealing with the byzantine pricing of some of the bigger platforms. But as we all know, the devil is never in the marketing page; it's in the implementation details and the throttling fine print.
Here's my initial autopsy after a month:
* **Ingestion & The Cardinality Trap:** The 10k *active* series limit for metrics is the first real constraint. It sounds like a lot until you realize how fast poorly optimized instrumentation or auto-discovered system metrics can chew through it. I hit a warning at 85% without even trying. The moment you brush against that limit, they start sampling. Not dropping, but sampling. Which means your graphs suddenly get... artistic. Less "observability" and more "impressionism."
* **Query Latency & The Dashboard Tax:** For basic PromQL queries on recent data, it's snappy. The second you try to build a dashboard that looks back more than a few hours, or aggregates across a few services, the lag becomes palpable. It's the classic "free tier" experience: you're not paying with money, you're paying with your time, waiting for panels to render. Alert evaluations seem to have higher priority, but I've seen delays of a minute or more on rule executions during what I assume are their internal load spikes.
* **The Logs Situation:** 50GB is generous, but the ingestion pipeline has a sort of soft, polite queuing behavior under burst. You won't lose logs, but they might take a scenic route, arriving 90 seconds later than their metric counterparts. This makes correlated troubleshooting a game of "wait for it." Retention is fixed at 30 days for the free tier, which is fine, but the query interface feels like it's on a diet compared to the paid query engines.
* **Alerting Reliability:** This is the part that makes me nervous. I've had a couple of test alerts fire 5-7 minutes after the threshold was crossed. For a true production-critical alert, that's an eternity. The alerting itself is reliable in that it *will* fire, but the "when" feels non-deterministic. I would not trust this for anything requiring SLA-level response times.
My overall take: it's a fantastic tool for a hobby project, a proof-of-concept, or maybe staging. But for "small production load"? That depends entirely on your definition of "production." If you need deterministic performance, guaranteed ingestion, and timely alerts, the free tier will feel like a straitjacket pretty quickly. It's a classic gateway drug—competent enough to get you hooked, then you'll start feeling the pain points exactly where they want you to: scale and reliability. So, who else has tried to live in this particular free-tier cage? Did you find the limits faster than I did, or am I just being my usual cynical self?