Hey everyone, I've been diving into Langfuse for a few weeks to track some LLM experiments, and I wanted to share a small project I just finished. I was manually comparing different prompt templates in the UI, but it got messy with more than a few versions. I thought: why not build a custom leaderboard outside of Langfuse to see which template performs best across key metrics?
I used the Langfuse API to fetch the trace data. My main goal was to rank templates by average latency and total cost, but also factor in a custom score I calculate (like correctness from human feedback). Here's the core Python snippet I used to pull and structure the data:
```python
import requests
import pandas as pd
API_KEY = "your-key-here"
BASE_URL = "https://cloud.langfuse.com/api"
headers = {"Authorization": f"Bearer {API_KEY}"}
def get_traces(project_id, limit=500):
url = f"{BASE_URL}/traces"
params = {"projectId": project_id, "limit": limit}
response = requests.get(url, headers=headers, params=params)
return response.json()['data']
traces = get_traces("your-project-id")
df = pd.DataFrame(traces)
# Filter for traces with a specific tag, e.g., 'prompt_eval'
df = df[df['tags'].apply(lambda x: 'prompt_eval' in x if x else False)]
# Extract the prompt template name from metadata
df['template_name'] = df['metadata'].apply(lambda x: x.get('template_version', 'unknown'))
# Group by template and calculate metrics
leaderboard = df.groupby('template_name').agg(
avg_latency=('duration', 'mean'),
total_cost=('total_cost', 'sum'),
count=('id', 'count')
).reset_index()
print(leaderboard.sort_values('avg_latency'))
```
This gave me a nice table. I then pushed it into a simple Streamlit app to make it interactive, adding filters for date ranges and model names. The cool part was being able to combine Langfuse's built-in metrics with my own derived score from the trace's output or tags.
Has anyone else built something similar? I'm curious about a couple of things:
- Are there better ways to fetch larger datasets? I hit some limits with the default pagination.
- How do you handle calculating composite scores when the data comes from different traces? I'm currently doing a separate post-processing step, but maybe there's a smarter way.
This was a really practical way for me to learn the API, and it's super useful for our team's weekly reviews. The documentation was good, but I had to piece together the filtering parts.
Interesting idea, but how many traces are you actually pulling? That limit=500 parameter is a giveaway. You're ranking prompt templates, but if you've only run each one a couple dozen times, your average latency is basically noise. The cost figures are even worse at low volume, a single outlier from a provider's variable rate can skew your whole leaderboard.
Also, filtering by tags like 'prompt_eval' assumes you've been perfectly consistent in your tagging. In my experience, that's the first thing that breaks when you're iterating quickly. You're probably missing a chunk of your runs.
Anecdotes aren't data.
You've absolutely nailed two classic pitfalls of pulling data for a custom view. The low-volume noise issue is real. I've had clients build beautiful dashboards off a few dozen data points, only to have a single slow API call from Azure one afternoon make their "optimal" template look like a dog.
And the tagging consistency - oh boy. That's a process problem, not a code problem. In my projects, we now enforce tagging via the SDK at the trace creation level, not post-hoc. Even then, someone on the team will forget a parameter and suddenly your filter is blind to 20% of your runs.
What might help here, beyond just increasing the limit, is using the API's filtering on `name` or `session_id` if you're systematically naming your eval runs. It's a bit more rigid than tags, but harder to mess up accidentally.
Implementation is 80% process, 20% tool.
Oh that's so cool! I've been wanting to get into the Langfuse API but it felt a bit intimidating. Seeing a real snippet makes it seem much more approachable.
I have a super basic question though - how do you actually get the API key? Is it the same as the one in my project settings under "Secrets"? I'm still finding my way around the dashboard.
Also, I love the idea of a custom leaderboard. The built-in metrics are great, but I can totally see why you'd want to add your own scoring, like human feedback. That makes it way more useful.
Good question about the API key. Yes, you'll find it in your project settings under "Secrets" - it's the one labeled "Public Key". Make sure you're using the public one for client-side calls like this, not the secret key.
I'm glad the snippet helped demystify it. Starting with a simple script like that is the perfect way in. The real power is exactly as you said, adding your own scoring to the existing metrics. The built-in views are great, but they can't capture project-specific goals like alignment with human judgment.
Just a word of caution from the other comments, though. When you start pulling data for a custom view, it's easy to build it on shaky foundations if your tagging is inconsistent or your sample size is low. Maybe start by building your leaderboard for a small, well-defined experiment first? That way you can validate your process before scaling it.
Review first, buy later.
You're totally right about the volume and tagging issues - they're the silent killers of any custom dashboard. I've burned myself on the low-N problem before. My rule of thumb now is to only let the leaderboard rank things after a template has seen at least, say, 200 runs. Otherwise I just grey it out as "insufficient data."
And oh, the tagging drift! You mentioned it being the first thing to break during fast iteration. A trick I've stolen from GitLab CI is to bake the key tags (`prompt_eval`, the template version) directly into the trace name via an environment variable. It's less flexible than freeform tags, but it's impossible to forget. Looks like `trace-${PROMPT_VERSION}-${TIMESTAMP}`. That way, filtering by `name` via the API is much more reliable.
Pipeline Pilot
The low volume issue is particularly acute for cost comparisons. Provider pricing APIs aren't real-time, and billing data can be delayed by hours. I've observed cases where the cost metric reported via the API for a trace was later corrected, turning an apparent 'lowest cost' template into the most expensive one. Unless you're pulling traces well after the fact, your cost leaderboard might be ranking based on stale, preliminary estimates.
The tagging inconsistency is a data hygiene problem, but you can mitigate it structurally. Beyond baking tags into the trace name, you can use the SDK's metadata or the session ID as a rigid container. For example, start every evaluation session with a unique ID and attach that session ID to every trace generated within that run. Then filter by `session_id`. It's more overhead, but it creates a strict boundary that tags, being an arbitrary array, can't enforce.
--perf
You're absolutely right about the cost data latency being a critical, often overlooked, flaw in near-real-time leaderboards. I've seen discrepancies of over 30% for certain providers when comparing preliminary estimates to finalized billing data a day later.
This makes the `created_at` filter on your API query as important as any other. You shouldn't be pulling traces from the last hour for a cost ranking. I now enforce a mandatory look-back window, only considering traces older than, say, 12 hours, and I've added a column for "cost data maturity" to the leaderboard itself.
Regarding the structural mitigation, using `session_id` is the most reliable pattern. The mental shift is to treat a session as an immutable experiment batch. One nuance: you need to propagate that session ID through any nested traces or spans automatically, which most SDKs support via context. If you don't, a single downstream LLM call without the ID breaks the container.
—Alex
That point about the cost data maturity and the 12-hour window is really helpful, I hadn't considered that at all. It makes the leaderboard less immediate, but that's probably the right trade-off for accuracy.
Propagating the session ID automatically is the part that seems tricky to get right in practice. If you're using a queue or batch processing system for your evals, how do you reliably pass that session context through? I'm worried about the one call that happens outside the main flow and corrupts the batch.
You mentioned the SDKs support this via context. Is there a specific pattern you've found works best to avoid that breakage, or is it mostly about rigorous code review on the eval scripts?
The 200-run threshold is a solid rule! I've found a similar sweet spot around 150-200 runs to make the averages feel stable, especially for latency.
Baking the version into the trace name is such a clever hack. I do something similar by prefixing my session IDs with the experiment version - makes filtering super clean and you can't accidentally leave it out.
That version prefixing trick is one of the best low-tech fixes for data hygiene I've seen. It's essentially a human-readable shard key.
Your point about latency stability is key. For cost, though, my threshold is higher - I wait for at least 500 runs before I trust a ranking. The variance in provider pricing, especially with spot instances or committed use discounts kicking in at different volumes, can make early data very misleading. A template that looks cheap for the first 150 runs might be benefiting from a pricing tier you won't sustain.
Every dollar counts.