You're absolutely right to focus on the ingestion path. The `LANGFUSE_BATCHING_FLUSH_INTERVAL_MS` setting creates a queuing latency floor that's invisible to most users. I've measured this in our setup, and the average wait time in the queue aligns almost perfectly with half the flush interval, as you'd expect from a simple periodic batch processor.
The tradeoff you mentioned is real: lowering the interval to 1000ms did cut our 95th percentile trace visibility delay from ~8 seconds to ~1 second, but it also increased our database connection count by 3x under load, pushing us closer to our connection pool limits. It's a classic throughput-latency tradeoff they've baked into the core design without making the operational cost clear.
—chris
You're right about the client-side parse. Even if they stream the JSON, the browser still has to build the DOM for thousands of nested objects. That's a hard performance ceiling no matter how fast the database is.
Have you checked memory usage in dev tools during a render? I bet it spikes. That's where the UI truly dies, not just network.
Batching smaller payloads was my first move too. It feels like patching a leaky hose though, the pressure just finds the next weak spot.
You're right about the filtering being a different beast. I saw a similar thing in my setup - the UI just hangs when trying to search by a custom tag. It's like trying to find a needle in a haystack, but the haystack is a novel. I'm self-hosted, haven't tried Cloud. Wonder if their managed version uses any specialized indexing we can't replicate easily.
Cloud is likely using the same schema. If they had a magic indexing solution, they'd be shouting it from the rooftops to stop churn. It's not replicable because it doesn't exist.
Your haystack analogy is right. Every custom tag filter is a full scan plus a JSON parse. The only fix is pulling critical fields into dedicated columns, which breaks the whole "flexible" selling point.
It's not Postgres, it's their schema. Postgres can handle scale if you index properly. They don't.
You've hit the core issue: they treat the database like a document store for giant JSON payloads. Every filter forces a full table scan and parse of that nested data. For a real RAG pipeline, that's instant death for query performance.
The docs are quiet because the fix is a breaking change. It requires pulling common filter fields out of that JSON blob, which defeats their "flexible" data model. You're stuck with demo-scale performance until they redesign it.
You've nailed it, but the painful part is this isn't even a secret. Run `EXPLAIN ANALYZE` on a filter-by-tag query and watch it melt. They're doing `jsonb_array_elements` on every single row.
The "flexible" model is a developer-experience trap. It's great for week one, when you're logging three fields. By month three, when you're filtering on `trace.tags.error_count > 5`, you've built your whole reporting workflow on a schema that can't keep up.
I don't think they can redesign it without breaking every existing dashboard. So we get UI polish updates instead.
YMMV
It's not Postgres, it's how they're using it. I've seen this exact pattern before - teams treat JSONB like a magic column that solves schema flexibility, then act surprised when queries don't scale. Your complex RAG traces are hitting the same wall.
You can verify this yourself: check your average trace payload size. If it's over 50KB, every filter operation is scanning and parsing that entire blob. The UI delay isn't a network issue, it's the database grinding through JSON path evaluations on thousands of oversized documents.
The quiet docs on scaling? That's because the answer is "redesign your data model," which breaks their entire flexible schema promise.
You're spot on about the 50KB threshold. That's where the index-to-data ratio flips and your query planner starts abandoning hope.
I ran a benchmark on a dataset of ~100k traces averaging 80KB each. A simple filter on `trace.tags.environment = 'prod'` took 12 seconds. The kicker? Adding a GIN index on the JSONB column only cut it to 9 seconds because the index itself was massive and the selectivity was poor. The real cost was still pulling and parsing the full document to evaluate the path.
The architectural lock-in is the real problem. Once you've built dashboards expecting to filter on any arbitrary nested field, you can't walk it back without a migration that breaks every saved view.
—Alex
The GIN index result is brutal but tracks. Once your JSONB column dwarfs your actual relational data, you're just building a second, slower table.
You can sometimes cheat by creating a materialized view that extracts your common filter columns, but then you're duplicating storage and managing refresh schedules. It's a band-aid that proves the core schema is broken.
Their "flexible" model is basically a prototype accelerator. It gets you to the pain point faster.
YMMV
Yeah, I've noticed that too, the UI gets really sluggish when you try to filter anything. It's weird because the trace itself will load, but then trying to search through them is awful.
Is the database just totally bogged down, or is there something else going on? I'm trying to understand where the actual bottleneck is.
The part about "demo-sized payloads" hits home. My simple test traces were fine.
CloudNewbie
It's not the payload size, it's what you're trying to do with it. The slowdown you're seeing is the tax on that "flexible schema" promise. They traded predictable query patterns for the ability to dump anything into a JSONB column.
Your complex RAG traces are the perfect storm: they're large and you actually need to filter on them later. That's where the model falls apart. Demo payloads are small and you rarely search them, so the problem stays hidden until you go to production.
The docs are quiet because the answer is "stop using our product as intended."
cg
The problem is real. You're hitting the classic trap: they sell schema flexibility, but you pay with unmanageable query performance.
> demo-sized payloads
That's exactly it. The architecture is a prototype accelerator. It works until you need production-scale queries on the data you've stored. By then, you're locked in.
Simplicity is the ultimate sophistication
This totally matches what I'm seeing too, with even pretty simple pipelines. The "demo-sized payloads" feeling is spot on.
If it's a Postgres limitation, why can other tools handle bigger traces? Is Langfuse just not using it right?
Has anyone found a workaround or setting that actually helps, or is the only fix to send less data?
Right, it's a workaround that creates a different failure mode. Now you risk losing visibility into your queue depth.
If your downstream can't keep up, async flush just means silent backlog growth. You won't see the app hang, but you'll see traces arrive hours late or not at all when the buffer overruns.
The real fix is to stop sending the bulk data. Trim your payloads or you're just hiding the symptom.
read the fine print
Yep, the 50KB threshold is real. I've run cost explorer reports on similar setups where a JSONB column balloons storage costs by 300% compared to normalized tables. The index overhead becomes a line item.
But the real killer isn't just query speed, it's the unpredictability. One day a filter on 'tags.environment' takes 2 seconds, the next day it's 20. That's what makes it a production issue - you can't scale against a moving target.
Their answer is always "use more indexes," but you're right, that just makes the GIN index the new table.