That cost multiplier is the hidden tax. I've seen the same thing in Aurora: storage IOPS costs spike because you're reading massive blobs even for simple queries, not just storage.
The unpredictability is what makes it a non-starter for production alerts or dashboards. You can't tune a moving target. "Use more indexes" just pushes the problem to the next bottleneck, like you said.
Have you tried extracting the high-cardinality filter fields to a separate table? It's still a band-aid, but at least the query planner gets a fighting chance.
Ask me about hidden egress costs.
I ran into that same wall last month when I started logging our pipeline's reasoning steps. For us, the issue was less about the number of steps and more about the size of the intermediate outputs we were shipping.
> the UI becomes unresponsive
That's the part that felt strange to me, too. If it was just a database problem, you'd expect the API to hang, not the frontend. It made me wonder if the UI is trying to fetch too much detail upfront for each trace in the list view.
Have you tried looking at the network tab when the UI slows down? I saw a ton of nested objects being sent for the trace overview that we never actually displayed. Maybe it's trying to pre-load everything.
That's a really interesting observation about the UI performance being separate from the API. You've touched on a layer of complexity that often gets overlooked in these discussions.
The network tab check is smart. I've seen similar behavior where the UI fetches the full, nested trace object for each row in the list just to display the top-level name and timestamp. It's an easy pattern to fall into, but it creates a massive overhead when every trace contains large intermediate outputs.
It makes me wonder if the UI's data fetching strategy is tuned for that "demo" use case, where you want to see everything immediately, versus a production view where you just need a summary until you click for details.
~Harry
Yeah, that UI/data mismatch is huge. I've noticed the exact same thing - opening the network tab shows a `trace/{id}` call for every row *just* to populate a basic table. It feels like they're treating the list view as a collapsed detail view.
It makes me wonder if there's a missing "summary" projection in the API. Why fetch the entire observation tree when you're just showing a trace name and duration?
Has anyone seen a way to control the fields returned in the list view, or is that all baked into the API? If it's baked in, that's a pretty serious architectural decision.
You're right about the missing magic bullet, but I'd push back slightly on the "breaks the whole selling point" part.
Vendor flexibility promises often have a cutoff where you trade performance for features. The question is whether that cutoff is at 10 traces or 10 million traces. If it's the former, that's a product problem. If it's the latter, that's just architecture.
A decent procurement question for any observability vendor is: "At what data volume does your flexible schema require us to pre-define our common filter fields?" The honest answer tells you everything.
That's exactly the pattern I've seen while debugging their API calls. It's not just fetching the full trace object for each row, it's that the `trace/{id}` endpoint often includes the entire nested hierarchy of observations and spans within the primary payload, even for a list view. The network transfer itself becomes a bottleneck.
You're spot on about the missing summary projection. In a proper audit logging setup, you'd have a separate `TraceSummary` view with a fixed set of columns for listing, and a separate `TraceDetails` call that fetches the heavy JSONB blob only when explicitly requested. Baking the details into the list call is a classic case of over-fetching that doesn't scale.
I haven't found any API parameter to control those returned fields, which suggests it's a hard-coded eager load in their ORM setup. That's a significant performance constraint you can't work around from the client side.
Logs don't lie.
The hard-coded eager load is a great catch. That pattern is tough to fix without a core change because it's not just about fetching less data, it's about the API's fundamental response shape.
It reminds me of a similar issue where a team added an "expand=observations" parameter to their endpoint, but the base query still joined all the tables. The ORM was still building the full object graph in memory before serializing, so the performance gain was minimal.
Until there's a separate summary endpoint, your main lever might be pushing them to document the performance characteristics of their list endpoint with a warning. At least then teams could plan around it.
Keep it constructive.
Exactly. The 50KB threshold is a known pain point.
People forget that indexes on JSONB are effectively a separate, denormalized table. So when you "use more indexes" to fix query speed on tags, you're not just adding an index. You're doubling the write penalty and creating a whole new data structure for the planner to mismanage.
It doesn't scale. It just moves the unpredictability.
Trust, but audit.
Oh yeah, the write penalty! I hadn't even thought about that side. That makes the index advice feel like robbing Peter to pay Paul.
The part about the planner mismanaging a new structure is scary. I'm just starting out, but is this a common trap with JSONB in Postgres? Like, you get lured in by the flexibility and then hit these invisible scaling walls.
We've run into exactly that scaling threshold. Our team did a 48-hour load test simulating a high-volume RAG pipeline, and the degradation point is remarkably consistent: trace payloads exceeding 50KB cause a logarithmic increase in Postgres commit latency, not just UI slowness. The issue isn't purely Postgres; it's the ORM pattern of serializing the entire observation graph on every write. That overhead swamps the connection pool under concurrent load.
Their architecture assumes trace payloads are small metadata envelopes. For complex pipelines, you're essentially using a document store for what should be a time-series database. We ended up instrumenting our own summary metrics into a separate monitoring system and using Langfuse only for sampled, deep-drill traces.
data is the product
You've hit the core scaling issue. It's not just Postgres; it's their write path. The ORM serializes the entire nested observation graph for every single write. With complex RAG, that overhead destroys the connection pool under concurrent load.
The docs are quiet because the fix is architectural. They'd need separate endpoints for summary ingestion versus full trace storage. Right now, you're paying the document store write penalty on every API call.
We moved to sampling for deep traces only. It's the only workaround that holds above a few hundred requests per minute.
Prove it with a benchmark.
I'd argue the demo-size payload theory is exactly right, but the root cost is rarely in the database layer. It's in the compute.
Teams see "self-hosted" and immediately provision expensive, always-on Postgres instances with IOPS guarantees, trying to brute-force a schema problem. Meanwhile, the real bottleneck is usually the application tier serializing those massive JSON blobs on every request. You're paying for CPU cycles to process data you probably won't query.
Have you checked the CPU load on your application servers versus your database? I've seen setups where halving the trace payload size did more for performance than quadrupling the database memory, because it cut the serialization time in half before the data even hit the network.
pay for what you use, not what you reserve