Hi everyone, I'm pretty new to all this and I think I've already made a classic blunder. I just started at a small startup, and my first task was to set up observability for our new web app. I read all the docs about "instrumenting your services" and thought, "more data is better, right?" So I added the agent to everything, used the default configurations, and turned on all the metrics, traces, and logs I could find.
Now we got our first bill from our observability platform, and it's... shocking. It's way over budget. I'm really embarrassed and my manager is asking questions. I clearly didn't understand how ingestion pricing works.
Could someone help me understand where I probably went wrong? I'm looking for basic advice on what to turn off or tune first. Things like:
- Are there specific, verbose logs that are common culprits?
- Should I be sampling traces instead of sending every single one?
- What are "high cardinality" metrics and how do I spot them in my setup?
I'm not sure which platform specifics matter here (we're using a popular cloud-native one), but I just need to stop the bleeding and learn how to do this properly. Any step-by-step guidance would be a lifesaver. Thank you so much in advance.
> "more data is better, right?"
That's the expensive assumption right there. Default configurations are designed to show off a platform's capabilities, not to be cost-effective. Your bill is huge because you're paying to ingest and index a firehose of data you'll never query.
Start with logs. Look for any debug-level logging, especially in frameworks. That's pure waste. Then, check for full request/response bodies being logged in web servers or API gateways; those are volume monsters.
On traces, yes, you absolutely need to sample. Start with a head-based sampler at a low rate, like 10%. Sending every single trace is for debugging a specific issue, not for continuous operation.
High cardinality metrics are your third leak. Look for any metric that uses a unique ID, email, or request path as a tag. Each unique combination creates a new time series. You'll see them as metrics with exploding dimensions in your UI. Prune those tags aggressively.
Been there, migrated that
Yes, starting with logs is the right move. That debug-level output from frameworks can just be turned off, it's never useful after development. But I'd be careful about sampling traces at 10% right away. If your traffic is low, you might miss the one weird error pattern. Maybe start higher, like 50%, and dial it down once you're sure you're catching the outliers.
Welcome to the club, that first bill is a rite of passage! The others gave good advice, but let me add something from the trenches: before you just turn things off, see if your platform has a usage explorer. It'll show you exactly which services or log attributes are generating the most volume. You'll likely find one or two noisy offenders you didn't expect.
I'd actually prioritize high-cardinality metrics next, even before touching trace sampling. A single metric with a user_id tag on it can explode your count faster than logs. Look for any tags that have unique values like IDs, emails, or full URLs. Change those to bounded values, like HTTP status code buckets or endpoint groupings.
Finally, don't just dial sampling down. Use a *tail-based* sampler if your platform supports it. It'll sample 100% of errors and high-latency traces, but drop the boring, successful ones. That way you keep the signal without the cost.
Ship fast, measure faster.
First, take a breath. We've all had a version of this moment, and you're doing the right thing by asking for help.
> I'm not sure which platform specifics matter
They do, but the principles are universal. Since you asked for a step-by-step, here's what I'd do tomorrow morning:
1. Log into your platform and find the "usage" or "ingestion" dashboard. Look for the biggest spikes in volume over the last 24 hours. That's your top offender.
2. For logs, search for the word "debug." If you see a lot, your framework is likely dumping stack traces or verbose internal states. Turn that log level to "info" or "warn" in your app configs. That's your quickest win.
3. For traces, you absolutely need sampling. Set up a head-based sampler at, say, 25% for now. That instantly cuts 75% of your tracing cost. You can refine it later.
High cardinality metrics are tricky to spot. Look for any metric with a tag that could have thousands of different values (like `user_id`, `request_path="/users/12345/profile"`, or `email`). If you find one, you need to change that tag to something bounded, like grouping status codes or using a generic endpoint name.
Your main job now is to reduce volume, not achieve perfection. Make those three changes first, monitor the usage dashboard for a day, and then you can tackle the next noisy thing.
ship early, test often