Skip to content
Notifications
Clear all

Step-by-step: Adding cost tracking to every LLM call in your LangChain application.

5 Posts
5 Users
0 Reactions
4 Views
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
Topic starter   [#28394]

Everyone's talking about building with LangChain, but I haven't seen a single post that mentions the first thing you should actually do: instrument the money drain. You're probably burning through API credits right now and have no idea which chain, which agent, which pointless "creative" rewrite step is responsible. The default callbacks are useless for this.

Forget the fancy tracing UIs for a second. You need a simple, brutal ledger. Here's the pragmatic approach I use. It's not pretty, but it tells you exactly where each cent goes.

First, you need a callback handler that intercepts every LLM call. The key is to hook into `on_llm_start` or `on_llm_end`, parse the token usage from the response, and apply the pricing for that specific model. Don't assume GPT-4 prices; Azure, Anthropic, and even different GPT-4 context windows have different costs. You must map the model name to a rate.

The core of it is a dictionary mapping model identifiers to cost per 1K tokens for input and output. Then, in your callback, you do the math: `(completion_tokens * output_cost / 1000) + (prompt_tokens * input_cost / 1000)`. Log this with a timestamp, the model name, and a context tag (like the chain name). Push it to a list, a thread-safe queue, or better yet, directly to a cheap database table. I use a Postgres table with columns for timestamp, model, prompt_tokens, completion_tokens, total_cost, and a metadata JSON field for the call context.

The biggest pitfall? Sample size. If you only run this in development with three queries, your "cost per chain run" metric is meaningless. You need to run this in production for a representative period to see the real distribution. Also, watch out for streaming responses; token usage reporting can be different.

Without this, you're flying blind. You'll optimize for "cleverness" instead of cost, and you'll get a nasty surprise when the bill comes. Implement this before you even think about moving to production.


Anecdotes aren't data.


   
Quote
(@alexgarcia)
Honorable Member
Joined: 2 months ago
Posts: 496
 

Spot on about mapping model identifiers. It gets messy fast when a chain silently switches from gpt-4-turbo to gpt-4o mid-stream because of a fallback setting.

One caveat: your cost dictionary needs constant maintenance. Providers change prices quietly, and new model versions roll out weekly. I've seen teams use stale prices for months. Maybe suggest a simple version check or pulling from a curated source?

The context tag for the chain or agent is the real killer feature though. Without that, you're just staring at a bill with no idea how to reduce it.



   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

The price dictionary maintenance problem is a significant one, and it's why I moved to a two-tiered lookup system. I keep a local fallback dictionary, but the primary source is a small internal API that fetches and caches the latest pricing from each provider's official pricing page daily using a simple scraper. This decouples the application code from the constantly shifting rate cards.

Even with that, the model identifier mapping is only half the battle. You need to also capture the *provider* from the LLM call's `run_id` or metadata to correctly attribute costs when using wrappers like LangServe or when multiple Azure endpoints with different billing structures are in play. Without that, you're just tracking tokens without knowing which invoice they'll appear on.

A final note on the math: remember to account for per-request costs if you're using certain Azure models or Claude with vision, where the token-based formula alone underestimates the bill.


Data over dogma


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

Your two-tiered system is a smart move for keeping prices current. The daily scrape from official sources is especially good for compliance, since it creates an auditable trail of the rates you assumed at billing time.

That point about capturing the provider from metadata is crucial. In a multi-vendor setup, I've seen teams correctly track tokens but then lump all Azure OpenAI costs together, missing the different markups per department or project baked into each endpoint's contract.

You're right about per-request costs. Those flat fees can quietly double the expected cost for high-volume, low-token operations, and they're easy to overlook if you're only counting tokens.


Review first, buy later.


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

The compliance angle is the only reason our finance department signed off on the custom scraper. They don't care about my pipeline latency, but they love that timestamped CSV it dumps.

You hit the real pain with the per-request costs. The official pricing pages often bury that detail in a footnote. Our scraper had to be taught to look for phrases like "per 1K input tokens, plus $0.002 per request" which, for a high-throughput summarization job, made the request fee 40% of the bill. If you're only logging the token counts from the LangChain callback, you're blind to it.

The departmental markups on Azure endpoints are a whole other accounting nightmare. We ended up tagging each call with a `project_id` pulled from a config map that ties the specific LLM client instance to a cost center. Without that, you're just moving the attribution problem one layer up.



   
ReplyQuote