Skip to content
Notifications
Clear all

Step-by-step: Adding a custom monitoring dashboard with Prometheus metrics.

3 Posts
3 Users
0 Reactions
0 Views
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
Topic starter   [#29324]

Hi everyone. I've seen a few threads recently about monitoring LangGraph applications in production, especially around tracking custom business logic and workflow-specific states. While the built-in LangSmith integration is fantastic for tracing, sometimes you need to integrate directly with your existing observability stack.

I wanted to share a practical approach I've been using to expose Prometheus metrics from a LangGraph workflow. The goal was to track things like the number of times a specific tool was called within a cycle, the distribution of decision paths taken by an agent, or the duration of specific subgraph executions. Here's a simplified breakdown of how I did it.

First, I created a simple metrics service module that initializes and holds the Prometheus metrics I care about. I used the `prometheus_client` Python library. I defined counters for tool calls and a histogram for step execution times. Then, in my graph's nodes, I added calls to increment these counters or observe the durations. The key is to instrument the actual functions the nodes call, not the graph structure itself.

For example, in a tool-calling node, I wrapped the tool execution logic. After the tool runs successfully, I'd call something like `metrics.tool_calls_counter.labels(tool_name='web_search').inc()`. To make the metrics available for Prometheus to scrape, I ran a simple HTTP server exposing the `/metrics` endpoint on a separate port from my main application.

The main pitfall to avoid is letting the metrics instrumentation significantly alter your graph's logic or error handling. Keep the metric calls as side effects that don't affect the state. This approach has given me a much clearer, unified view of my graph's behavior alongside my other services in Grafana. Has anyone else tried similar custom instrumentation? I'm curious about other patterns for tracking state graph-specific metrics like loop iterations or conditional branch ratios.

—HR


—HR


   
Quote
(@dianar)
Honorable Member
Joined: 2 months ago
Posts: 487
 

Instrumenting at the function level is the right call. It gives you clean, actionable metrics tied to business logic.

Did you run into cardinality issues labeling those counters? If you're using tool names *and* workflow states as labels, it can explode quickly. We had to add a static allow-list for valid label combinations.

Also, make sure you're exposing the metrics endpoint on a different port than your app. You don't want that scrape endpoint being load-balanced or exposed publicly.


Five nines? Prove it.


   
ReplyQuote
(@harperl)
Estimable Member
Joined: 3 months ago
Posts: 127
 

Oh, good point about the port. I was planning to put it on the main app path. Why exactly shouldn't it be load-balanced? Is it just a performance thing for the scraper, or a security risk too?

Also, what do you mean by a static allow-list for labels? Like you only let certain tool/state combos be recorded?


Ask me in a year


   
ReplyQuote