Skip to content
Notifications
Clear all

Walkthrough: Building a CI/CD gate using LangSmith regression test scores.

1 Posts
1 Users
0 Reactions
17 Views
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
Topic starter   [#9192]

Just built a CI/CD quality gate using LangSmith regression test scores. It's a solid way to catch LLM degradation before it hits production, but I'm curious if others have measured the actual cost-benefit.

Here’s the core workflow we implemented:

* Our CI pipeline (GitHub Actions) triggers a LangSmith dataset run via API after any model/prompt change.
* We fetch the aggregate scores (e.g., correctness, faithfulness) and compare them against a threshold defined in the pipeline config.
* If scores dip below the threshold, the pipeline fails. No deployment.
* We also track score history in a simple dashboard for trend analysis.

Key questions for the community:
* What's a realistic threshold for, say, "correctness" in a production app? 0.95? 0.98?
* How do you handle the cost of running these regression tests on every commit? Is the ROI there compared to the risk of a bad deploy?
* Any pitfalls with using the aggregate scores as a single gate? Found we needed to add a check for extreme outliers on individual test cases too.

—CR


Ask me about hidden egress costs.


   
Quote