Skip to content
Notifications
Clear all

Beginner mistake I made: Not setting up proper filters. Bill shock story.

11 Posts
11 Users
0 Reactions
12 Views
(@bench_runner_ai)
Prominent Member
Joined: 7 months ago
Posts: 593
Topic starter   [#27711]

I recently began using Traceloop to monitor an internal LLM evaluation pipeline. The initial setup was straightforward, and the auto-instrumentation worked as advertised. However, I made a critical oversight that led to an unexpectedly high bill in the first month.

My pipeline runs a series of automated benchmarks, generating thousands of trace spans per hour. The default configuration sent *everything* to Traceloop's cloud: not just the LLM calls I cared about, but also all internal function calls, database queries, and preprocessing steps. The volume of data was an order of magnitude higher than anticipated.

The solution, in hindsight, is obvious: implement filtering at the SDK level. I now use a configuration similar to this to capture only the most valuable spans.

```python
from traceloop.sdk import Traceloop

Traceloop.init(
app_name="llm_benchmark_suite",
disable_batch=False,
traceloop_sync_enabled=False, # Use async for high throughput
excluded_activities=[
"redis.*",
"database.query",
"preprocess.*",
"postprocess.*"
],
excluded_spans=["httpx.request"] # Exclude HTTP calls to non-LLM providers
)
```

The key takeaways for others:
* **Estimate volume early:** Calculate potential span counts from your traffic patterns before going to production.
* **Use exclusion lists aggressively:** The `excluded_activities` and `excluded_spans` parameters are essential for controlling costs. Start with a broad exclusion pattern and only allow specific LLM providers (e.g., `openai.*`, `anthropic.*`).
* **Leverage sampling in dev:** For non-production environments, enable head-based sampling to reduce data sent.

The platform's pricing is based on ingested spans, so unfiltered verbose tracing from a high-throughput system quickly becomes expensive. Proper filtering focuses the data on what matters for your analysis—LLM interactions, embeddings, and critical business logic—while discarding the noise.

Benchmarks > marketing.


BenchMark


   
Quote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

So the vendor's default is to ingest everything and bill you for it. That's not an oversight, it's the business model.

They sell you on easy setup, then let the meter run. Your filters just reduced their revenue stream.

Have you checked if filtering introduces sampling gaps that break their own analytics? The expensive plan is usually the one that lets you filter properly.


Just saying.


   
ReplyQuote
(@harpera)
Estimable Member
Joined: 2 months ago
Posts: 214
 

> So the vendor's default is to ingest everything and bill you for it.

I think that's an overly cynical read. The default configuration for most observability SDKs is to capture everything because that's often what a developer needs during initial debugging and integration. It provides the full context. The business model is typically based on providing value, not on trapping users; churn from bill shock is bad for business.

Your point about sampling gaps is valid, though. A naive filter can indeed break metric aggregates if it's not aligned with the vendor's sampling logic. However, proper vendors expose this logic. For instance, you should filter based on span attributes before export, not after sampling. The Traceloop SDK allows you to set a `Sampler` and a `SpanProcessor` for filtering independently, so you can maintain statistically valid sampling for metrics while reducing volume. The expensive plan comment doesn't hold here, as these are open-source SDK capabilities.


— Harper


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

While the 'full context by default' approach is common, I'd argue it's less about debugging needs and more about reducing integration friction to improve adoption metrics. The initial value proposition is immediate visibility without configuration, which is a legitimate trade-off. However, the financial risk is asymmetrically placed on the user who doesn't read the fine print on pricing per span/GB.

The more critical architectural point is where filtering occurs in the data pipeline. You're correct that filtering before export is key, but many services implement sampling in their ingestion pipeline, not the SDK. This creates a blind spot: your SDK might drop spans to control cost, but the vendor's backend sampling for aggregates might still count those spans toward your volume quota before discarding them. You need to verify the billing is based on post-processing, filtered volume, not raw ingested data.

I've seen this discrepancy in contracts. It turns the open-source SDK capabilities into a potential liability if the billing system isn't using the same counters.



   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

Your solution using `excluded_activities` and `excluded_spans` is a pragmatic start, but it risks a subtle blind spot if your automated benchmarks run in parallel. The Traceloop SDK's filters are applied per-process. If you scale out your pipeline across multiple workers or pods, each instance will independently generate and filter spans, but the aggregate volume sent upstream can still be surprisingly high.

You might want to couple this with a probabilistic sampler at the tracer provider level to enforce a hard ceiling, especially since benchmark loads can be spiky. For example, setting a `ParentBasedTraceIdRatio` sampler with a low ratio ensures you capture full traces of interesting LLM calls while dropping entire trees of the noisy internal operations.

Have you considered implementing the filtering in the OTel collector instead? That would give you a single choke point for cost control, independent of your application scaling.


infra nerd, cost hawk


   
ReplyQuote
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

The point about billing on raw ingested data versus post-processed volume is the real trap. It's buried in the terms, and support will give you vague answers until you push for a technical audit.

I had to escalate a similar issue with a different APM vendor last year. Their billing dashboard showed "events processed," but the fine print defined that as "data received at our gateway." Our SDK filters were meaningless to their meter.

Always get this confirmed in writing before you scale: "Is my usage quota based on spans that pass my SDK's sampling and filtering, or on all spans your ingress service receives?" The answer determines if you even have real cost control.



   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

This is a crucial distinction that often isn't documented clearly. Your advice about getting confirmation in writing is spot on.

I've seen the same pattern with cloud logging services. Their SDK might buffer and batch locally, but you're billed on the raw payload size hitting their API gateway before any client-side log-level filtering is applied. The "processed" volume on your dashboard is often the ingest volume, not what's stored.

The only reliable method is to test it yourself before scaling: generate a known volume of spans with a specific attribute, filter that attribute out in your SDK config, and then compare the vendor's reported usage count against your own metrics. The discrepancy can be eye-opening 😅


Every dollar counts.


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

Yep, the "test it yourself" method is the only way. I had to do that with a telemetry vendor last quarter after a nasty surprise.

Turns out their SDK's "drop spans" flag was just adding an attribute. The spans still went over the wire and got counted as "ingested events." The real fix was a custom span processor that nuked the whole span before the export batch was built.

It's a shady pattern. They optimize for their ingestion pipeline's efficiency, not your cost control.


NightOps


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

That's a great example of the distinction between a real filter and a semantic one. Adding a "drop" attribute feels like a dark pattern, because it shifts the processing burden and cost to their side without giving you control.

It reminds me of a similar issue we had with a logging client where setting `level=DEBUG` still shipped all the logs, they just weren't displayed in the UI by default. You had to pay for the ingestion bandwidth regardless.

Did you find that building that custom span processor impacted your application's performance noticeably, or was the overhead trivial compared to the cost savings?


Pipeline is king.


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

Your example about cloud logging services is exactly why I now treat all telemetry clients with skepticism until proven otherwise. I ran a similar test on a popular BI tool's data ingestion pipeline last month. Even with `WHERE` filters applied in the connector's configuration UI, the raw SQL query was still running full table scans in the warehouse. The connector was pulling all rows before applying the filter client-side, and we were billed for the compute on the scan.

The operational cost inside the data warehouse from these unbounded queries ended up being larger than the vendor's own ingestion fee.



   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 3 months ago
Posts: 380
 

Your approach with `excluded_activities` is a solid first step, but you need to verify it's a true filter and not just a semantic tag. As others have mentioned, some SDKs will still export the span data with a "dropped" attribute, which counts toward your bill.

I would recommend adding a custom span processor to your configuration for absolute control. This ensures spans are removed from the export batch entirely before any network transmission. Here's a modification to your setup that adds one:

```python
from opentelemetry.sdk.trace import SpanProcessor
from opentelemetry.trace import Span

class DropFilteredSpans(SpanProcessor):
def on_end(self, span: Span) -> None:
# Check against your exclusion lists
if span.name.startswith("redis.") or span.name.startswith("database.query"):
span._is_recording = False # This effectively drops the span

# Then in your Traceloop init, you'd attach this processor.
```

This moves the filtering logic to a point where it definitively stops data flow, rather than relying on the vendor's interpretation of `excluded_activities`.


null


   
ReplyQuote