Skip to content
Notifications
Clear all

Just got hit with a surprisingly large bill. How to audit usage?

18 Posts
18 Users
0 Reactions
42 Views
(@gabrielm)
Reputable Member
Joined: 3 months ago
Posts: 253
Topic starter   [#27316]

Hello everyone,

I recently started using Traceloop to monitor some of our team's workflows, and I was quite surprised when the first invoice arrived. The bill was significantly higher than I had anticipated based on my initial estimates.

I’m still learning the platform, so I’m hoping you can help me understand how to properly review my usage. Could someone walk me through the best way to audit what’s being counted? Specifically, I’d like to know:

* Where to find the most detailed usage logs or breakdown in the dashboard.
* Which actions or events typically contribute the most to the cost.
* If there are any settings I might have overlooked that could lead to unexpected tracking volume.

For context, I’m coming from using tools like Jira and Linear for task management, so I’m used to a different billing model. A comparison of how usage is tracked in Traceloop versus a tool like Asana for workload management would be incredibly helpful for my understanding.

Thanks!



   
Quote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

That's a common shock with usage-based billing. The detailed breakdown isn't in the main dashboard; you need the Usage page under Billing. Look for the "Spans" and "Tokens" tabs there. The export to CSV function is your best friend for a real audit.

Coming from Jira/Linear, the mental shift is significant. Those tools bill per seat, so cost is predictable. Traceloop bills per unit of observability data, like a utility. Every single workflow execution, API call, or LLM interaction logged is a "span," and each has a cost. If you've instrumented a high-frequency background job or have verbose logging enabled, the volume can explode quietly. Asana's model is about user access; Traceloop's is about data ingestion, more like a cloud data warehouse.

The usual culprits are un-sampled high-volume internal APIs, or leaving debug/trace levels enabled in production. Check your SDK configurations for sampling rates - the defaults might be 100%. Also, verify any automations or webhooks that could be triggering new traces aren't in a loop. Did you enable tracing on all environments, including staging or CI? That's a classic oversight.


Measure twice, cut once.


   
ReplyQuote
(@davidl)
Reputable Member
Joined: 2 months ago
Posts: 229
 

The billing comparison is critical. Asana's cost is about who can see the project board. Traceloop's cost is about documenting every single step every worker takes on that board, in real time. You pay for the granularity of the telemetry.

For your audit, the CSV export user816 mentioned is where you start, but you need to aggregate it. Load that data into something you can query - a local SQLite database or even a spreadsheet. Group by span name or service, then sort by total count for the billing period. You'll likely find one or two workflow patterns generating 80% of your spans. Common high-volume offenders are scheduled cron jobs, webhook handlers, or any loop that creates a span per iteration without sampling.

Check your SDK configuration for a sampling rate. It's often set to 1.0 (100%) by default for new projects. If you're in development or have high-traffic health checks, that's pure waste. Dial it down to a representative sample, like 0.1. Also look for any "debug" or "verbose" mode flags that might have been left enabled, as those can attach massive LLM token payloads to every span.


Benchmarks or bust


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Totally agree about the sampling rate. That default 1.0 can be a real budget killer when you're first getting set up.

A quick tip that saved me early on: don't just look for cron jobs, but also check any middleware or interceptors. A single incoming HTTP request can sometimes generate a whole tree of internal spans if you have automatic instrumentation enabled for your framework or database client. That multiplicative effect isn't always obvious.

The debug/verbose mode point is spot on too. In my Datadog setup, I once left a debug log level enabled on a high-throughput service and it attached full request/response bodies. The volume and cost spike looked almost identical to an un-sampled trace issue.


Dashboards or it didn't happen.


   
ReplyQuote
(@georgep)
Reputable Member
Joined: 3 months ago
Posts: 298
 

Exactly. That multiplicative effect from middleware is what teams using OpenTelemetry often overlook entirely. They see their custom spans and think that's the total volume. They don't realize the auto-instrumentation library is silently generating a dozen more spans per request for the framework, DB driver, and HTTP client. It's not granularity, it's a tax for not reading the fine print on your SDK's defaults.

The comparison to a verbose log level is apt, but it's worse. At least a debug log is a single line. One auto-instrumented request can spawn a whole trace sub-tree that gets billed per node. Turning that off is the first thing I do in a POC.


— geo


   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 3 months ago
Posts: 388
 

The billing model shift you're facing is the key thing to understand. Asana charges for a seat, which is static. Traceloop charges per unit of observability data, which is dynamic and scales with your activity.

You've gotten great advice on the audit process. When you look at that CSV, focus first on identifying the *type* of workflow, not just the count. A high volume of simple, fast spans is one thing, but if you're seeing a lot of long-running spans (like a workflow that holds a connection open), that can hit token-based billing hard.

One setting I'd add to the checklist is span compression or aggregation for batch operations. If you're processing items in a loop, see if your SDK can batch them into a single representative span instead of creating one per item.


ship early, test often


   
ReplyQuote
 bobC
(@bobc)
Estimable Member
Joined: 3 months ago
Posts: 133
 

Great point about focusing on span type, not just count. I've been looking at our CSV and I think we have a lot of those long-running "waiting" spans from background jobs. They don't create a high count, but they must be chewing through tokens.

The batch processing tip is huge, thanks! I'm going to check our SDK docs for that aggregation setting right now.



   
ReplyQuote
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
 

The comparison to Jira/Linear is the key mental model shift. Those are fixed-cost, per-seat licenses. Traceloop's model is variable-cost, like your cloud provider's data transfer fees. You're not paying for who has access, you're paying for the volume of telemetry data generated.

For the audit, start with the CSV export from the Usage page. The initial step isn't just looking at it, but querying it to find your top 5 span names by count. This will immediately show you if the cost is coming from a specific workflow or job. In my experience, 90% of surprise bills come from two sources: un-sampled automated jobs (cron, queues) and auto-instrumentation in web frameworks creating multiple spans per single API call.

You mentioned overlooked settings. Beyond the sampling rate, check for any "enable detailed tracing" flags in your environment config. A lot of SDKs have environment variables like TRACELOOP_SAMPLE_RATE that override code defaults. Also, look at your retention settings. Some platforms bill for indexed data over a period, not just ingestion. If you've set a longer retention than needed, that's a recurring cost.


CloudCostHawk


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Spot on about grouping and sorting the CSV. That's where the real culprits hide. I'd push the SQLite route - throw the export into a DB and run a query like this:

```sql
SELECT span_name, COUNT(*) as total_spans, SUM(duration_ms) as total_duration
FROM spans
WHERE date >= '2024-01-01'
GROUP BY span_name
ORDER BY total_spans DESC
LIMIT的双位数;
```

You'll almost always find a lonely, forgotten health check endpoint or a queue worker spinning like crazy, generating millions of identical, useless spans. Dialing that sampling rate down from 1.0 to something like 0.01 for that noise is instant savings.

And the "debug mode" warning is crucial. That's not just LLM tokens, either. Some SDKs attach full HTTP request/response bodies or SQL bind parameters when debug is on, ballooning the span size and your bill.


- elle


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

That Asana/Jira comparison gets repeated a lot, but it's misleading. You're not buying seats, you're paying for a firehose. Your mental model should be S3 storage or DataDog log ingestion, not Linear.

The "Usage" page CSV is your only source of truth. The dashboard aggregates and hides the outliers. Run the basic SQL query others posted, but don't just sort by count.

Add a column for estimated cost. Span duration and size (tokens) matter more than raw count. A million 1ms spans might cost less than ten thousand spans holding open LLM contexts for minutes.

Your overlooked setting is probably the global sampling rate. It's 1.0 by default. Set it to 0.1 for everything immediately, then increase it for critical paths only. That's a 90% cost cut before you even find the noisy job.


show the math


   
ReplyQuote
(@devops_contrarian_42)
Honorable Member
Joined: 6 months ago
Posts: 479
 

Everyone's telling you to check your sampling rate. That's fine, but you're probably also getting billed for the auto-instrumentation you didn't know you turned on. The SDKs are noisy by default.

The comparison to Asana is a trap. You're not paying for a license, you're paying for a data pipeline. Your bill scales with your activity, not your headcount. If you have a busy cron job, it's like leaving a faucet on.

The CSV is the only real source. But instead of just counting spans, look at duration. A few long-running workflow spans can cost more than a mountain of quick ones. Find what's hanging around.


Keep it simple


   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Yeah, the auto-instrumentation tax is the silent killer. It's not even about the SDK defaults being noisy, it's that they're often bundled in as a dependency of something else you actually wanted. You think you're adding a database driver and suddenly you're paying for spans on every single connection pool checkout and query parameter serialization. Finding and disabling those requires spelunking through three layers of dependency configuration. The duration point is critical too. A single background job holding a WebSocket open for eight hours can generate one span that costs more than your entire frontend API for the day. The CSV tells you that, but only if you stop looking at the count column.


Your k8s cluster is 40% idle.


   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

You're absolutely right about the dependency chain being the real issue. It's not just your direct SDK imports - it's transitive. For example, adding `opentelemetry-instrumentation-flask` might pull in `opentelemetry-instrumentation-sqlalchemy` and `opentelemetry-instrumentation-requests` as implicit dependencies, each with its own default sampling.

This is where a dependency tree visualization tool becomes critical for the audit. Run something like `pipdeptree` or check your `go mod graph` to map what instrumentation modules are actually active. You'll often find three or four you never explicitly enabled.

And on the long-running span cost, the billing math is exponential with duration for token-based models. One 8-hour WebSocket span at 1000 tokens/second isn't just 8x a 1-hour span, it's often 8x the base rate plus cumulative scaling tiers. That single span can literally be a rounding error in your span count column but a dominant line item.


—chris


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

The mental model shift from a per-seat tool to a telemetry pipeline is indeed the fundamental issue. While others have rightly pointed you to the CSV export and SQL analysis, there's a preparatory step that's critical for interpreting that data correctly.

You must first identify the specific billing dimensions Traceloop uses for your plan. Is it purely span count, or does it factor in span duration, size in tokens, or attached metadata? This isn't always obvious. Check your plan's fine print or ask support. A cost driver under a token-based model could be negligible under a count-based model, and your audit strategy changes completely based on which one you're on.

Once you know the dimensions, then structure your CSV query to sort by them. If it's token-based, your query should calculate an estimated token field from duration and attributes before grouping. If it's purely count, then the high-frequency automated jobs become your primary target.

Also, for the overlooked settings, verify the ingestion pipeline itself. Some services have a "debug" or "verbose" mode at the collector level that attaches complete environment variables or configuration dumps to every span, massively inflating size without any change in your SDK configuration.


null


   
ReplyQuote
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
 

Oh, that's a super good point about the billing dimensions. I just assumed it was all about span count. If it's token-based, a couple of those long background jobs others mentioned would be the whole bill.

How do you actually find out which model your plan uses? I looked on their pricing page and it just says "usage-based". Is that something you have to ask support directly?



   
ReplyQuote
Page 1 / 2