Hey folks, been diving deep into Langfuse for tracing our LLM pipelines and overall it's been a game-changer for visibility. But I've hit a major performance snag that's making me reconsider for high-volume, data-heavy use cases.
When we log traces containing large payloads—think full PDF text chunks, lengthy tool call responses, or even sizeable embeddings arrays—the Langfuse UI becomes painfully sluggish. The trace page takes forever to load, scrolling is janky, and sometimes the browser tab just freezes. This happens both in their cloud and our self-hosted setup (using the official Docker compose).
Has anyone else experienced this? I'm trying to pinpoint if it's our implementation or a known limitation.
Here's a simplified version of what our trace creation looks like when the problem occurs:
```python
from langfuse import Langfuse
import json
langfuse = Langfuse()
# Simulating a large payload
large_payload = {"chunks": ["...very long text..."] * 100} # Large array
trace = langfuse.trace(
name="large_doc_processing",
input=large_payload # This seems to be the culprit
)
# Subsequent spans/generations also slow if they have big outputs
generation = trace.generation(
name="big_summary",
output={"summary": "...huge summary text..."}
)
```
**Observations so far:**
* The slowdown seems directly proportional to the size of the `input` or `output` fields logged.
* It's primarily a UI/query issue; the SDK itself logs fine.
* The trace list view is okay, but drilling into a single trace is where it dies.
Are there any known workarounds? Should we be truncating data before sending it to Langfuse, or is there a configuration tweak we're missing? The data is invaluable for debugging, but the performance trade-off is tough.
Would love to hear if others have benchmarks or solutions—maybe a different storage backend config?
--weaver
Oh yeah, we hit this exact wall. It's not just you.
The UI chokes on huge raw payloads, especially the ones rendered in those expandable JSON viewers. It tries to render everything at once. We started using the `metadata` field exclusively for anything bigger than a few KB, and kept `input`/`output` for lightweight, representative samples (like a chunk count or a hash). It's a workaround, but it made the UI usable again.
Have you checked your browser's network tab while loading a trace? You'll likely see it's downloading the entire massive payload as part of the trace fetch. That's the core bottleneck.
Sleep is for the weak
Good call on checking the network tab - seeing those multi-megabyte transfers really drives it home. The metadata field trick is clever for the UI, but then you lose the ability to search and filter on that content, right? We had the same dilemma.
What's your threshold for "few KB"? We settled on 10KB for input/output, anything bigger gets a JSON pointer in the output and the full data goes to S3 with a link in metadata. It's extra plumbing, but the UI stays snappy and we keep the data.
It's not marketing, it's logic.