I've been benchmarking several LLM orchestration frameworks, and LangGraph's stateful, graph-based workflows are excellent for complex, multi-step reasoning. However, I found its built-in observability a bit limited for production monitoring. You get basic tracing via LangSmith, but what about custom business metrics, alerting, or a dedicated dashboard?
I decided to integrate Prometheus, the open-source systems monitoring toolkit. This allows you to expose custom metrics (e.g., graph step duration, failure rates, token counts) and visualize them in a Grafana dashboard. Here's a step-by-step report on how I implemented it.
First, you'll need to instrument your `StateGraph`. The key is to wrap your node functions to collect timing and call counts. I used the official `prometheus_client` Python library.
```python
from prometheus_client import Counter, Histogram, generate_latest
from functools import wraps
import time
# Define custom metrics
GRAPH_STEP_DURATION = Histogram('langgraph_step_duration_seconds', 'Duration of graph node execution', ['node_name'])
GRAPH_STEP_CALLS = Counter('langgraph_step_calls_total', 'Total calls to graph node', ['node_name'])
GRAPH_ERRORS = Counter('langgraph_errors_total', 'Total graph execution errors', ['error_type'])
def monitor_step(node_name):
def decorator(func):
@wraps(func)
def wrapper(state):
GRAPH_STEP_CALLS.labels(node_name=node_name).inc()
start_time = time.time()
try:
result = func(state)
duration = time.time() - start_time
GRAPH_STEP_DURATION.labels(node_name=node_name).observe(duration)
return result
except Exception as e:
GRAPH_ERRORS.labels(error_type=type(e).__name__).inc()
raise
return wrapper
return decorator
# Apply to your nodes
@monitor_step("research_agent")
def research_node(state):
# Your node logic
return {"results": ...}
```
Next, expose the metrics via a simple HTTP endpoint (e.g., using FastAPI) for Prometheus to scrape:
```python
from fastapi import FastAPI, Response
app = FastAPI()
@app.get("/metrics")
def get_metrics():
return Response(generate_latest(), media_type="text/plain")
```
Finally, configure Prometheus to scrape your running LangGraph app and create Grafana dashboards. You can now track:
- P95 latency per node 🚀
- Error rate spikes
- Total graph invocations
- Custom business logic (e.g., "documents_retrieved")
The main pitfall? Ensure your metric labels are low-cardinality to avoid performance issues. This setup provides the granular, custom monitoring needed for serious load testing and cost attribution. Has anyone else implemented a similar monitoring layer? I'm curious about alternative approaches or if you've found valuable specific metrics to track.
garbage in, garbage out