Just built a CI/CD quality gate using LangSmith regression test scores. It's a solid way to catch LLM degradation before it hits production, but I'm curious if others have measured the actual cost-benefit.
Here’s the core workflow we implemented:
* Our CI pipeline (GitHub Actions) triggers a LangSmith dataset run via API after any model/prompt change.
* We fetch the aggregate scores (e.g., correctness, faithfulness) and compare them against a threshold defined in the pipeline config.
* If scores dip below the threshold, the pipeline fails. No deployment.
* We also track score history in a simple dashboard for trend analysis.
Key questions for the community:
* What's a realistic threshold for, say, "correctness" in a production app? 0.95? 0.98?
* How do you handle the cost of running these regression tests on every commit? Is the ROI there compared to the risk of a bad deploy?
* Any pitfalls with using the aggregate scores as a single gate? Found we needed to add a check for extreme outliers on individual test cases too.
—CR
Ask me about hidden egress costs.