Hey folks! I'm trying to set up Langfuse to trace my LlamaIndex RAG pipeline in a production-ish environment. My proof of concept works fine, but I'm worried about scaling.
I'm getting ready to handle maybe 50-100 requests per second. Has anyone run this combo at high volume? My main concerns are:
* Does the Langfuse SDK add a lot of latency?
* Are there any Terraform modules or AWS config tips (I'm on ECS Fargate) to make the observability backend (Postgres/ClickHouse) more robust?
* I'm tracing spans for each step (retrieval, synthesis). Will this generate too much data and blow up costs?
Here's my basic tracing setup right now:
```python
from llama_index.core import Settings
from llama_index.core.callbacks import CallbackManager, LlamaDebugHandler
from langfuse.llama_index import LlamaIndexCallbackHandler
langfuse_callback = LlamaIndexCallbackHandler()
Settings.callback_manager = CallbackManager([langfuse_callback, LlamaDebugHandler()])
```
Does this look right for production? Should I be sampling traces instead of capturing everything? 😅
Any war stories or config snippets would be super helpful.
We run about 40 req/s. The SDK latency is negligible if you run the Langfuse client async. The bigger issue is the volume of spans.
For your scale, you'll need to sample. We keep 100% of traces for error cases, but only 10% for successful requests. You can set this in the Langfuse callback init. Without sampling, the data cost is insane.
On AWS, run the Langfuse backend separate from your app, obviously. Use provisioned IOPS for Postgres if you go that route. Fargate is fine for the client.
Thanks, that's helpful. I hadn't considered sampling at the callback level, I was only thinking about doing it in the data pipeline itself.
How do you define an error case for the 100% sampling? Are you catching exceptions and setting a flag, or is it based on something like an HTTP status?
PipelinePadawan
Your setup's a good starting point, but you'll definitely want to move beyond capturing everything for 50-100 RPS. The callback handler you have is synchronous, which will add latency.
Switch to the async Langfuse client and set sampling directly in the callback initialization. Start with a low sample rate for successful traces and keep it at 100% for errors, which you can define based on your application's HTTP status codes or by catching specific exceptions in your pipeline logic.
For your AWS setup, the key is separating the observability backend. If you stick with the default Postgres, monitor your write throughput closely - that's usually the first bottleneck at your volume.
Keep it constructive.
Yeah, the async client switch is critical. I'd also add you shouldn't even wait for the Langfuse `flush()` in your request path. Toss it to a background thread and let it handle its own batching and retries.
> monitor your write throughput closely
This is the real talk. At 100 RPS with even minimal spans, you're hammering the DB. We had to move to ClickHouse for the Langfuse backend to keep ingestion costs sane. Postgres will buckle unless you over-provision massively, and even then the storage cost for traces adds up fast.
pipeline all the things
That ClickHouse point is really interesting - I hadn't considered that the database choice would be so critical at high volume. How was the migration from Postgres? Did you lose any query flexibility or features by moving to ClickHouse for the trace storage?
Oh wow, this is exactly the kind of setup I was just looking at! I'm also building a RAG pipeline and had the same question about sampling.
Your callback setup looks like mine, but everyone here is saying it adds latency if it's synchronous. How do you even switch it to async? Is it just a different import, or do you have to rewrite how the callback manager is called in your queries?
The cost warning is scary. I was planning to trace everything to debug, but at 50 requests a second I can see how that gets huge fast.