It's both. Initial design is a day. The iteration loop is weeks. Stabilizing scores means running hundreds of evaluations against your gold set, tweaking prompt wording, weights, thresholds, and then watching for drift.
The only shortcut is a gold dataset with clear, unambiguous edge cases. If your ground truth is fuzzy, your metric will be noisy. For sentiment, we had to define "consistency" as "no polarity reversal within a three-sentence sliding window" before we could even start.
>Defining a custom metric is literally just writing a Python class with an `evaluate` method.
That simplicity accelerates initial development, but it shifts the burden of operational rigor onto your team. In a CI context, you'll need to wrap those classes with telemetry for latency percentiles and score stability across runs, else you risk metrics that degrade silently.
For brand voice adherence, an embedding similarity score against a versioned corpus is sound, but have you established a performance baseline for the embedding model itself? Drift in that component could invalidate your evaluations without clear attribution in the dashboards you'd build.
Exactly. That "ground truth fuzziness" is the silent killer for any CI/CD integration. You can have the cleanest class definition, but if your gold set doesn't capture edge cases well, you'll be chasing variance forever.
Your three-sentence window definition is a great example of making it concrete. We ran into similar issues with "tone consistency" and had to literally define it as "no shift into a formal/informal pronoun set within a single answer." That specificity cut our iteration loop in half.
How are you versioning that gold dataset, by the way? We've started treating ours like a test suite, tagging commits where we expand it, but I'm curious if others have a smoother method.
Webhooks or bust.
That pytest integration is the killer feature for CI. But watch the cold starts. If your DeepEval tests spin up an LLM for every custom metric, your pipeline's execution time and cost can balloon fast.
For CI, we ended up running evaluations in batches offline and just checking a versioned score file. The pytest integration was still useful for local dev, but we had to break the pattern for the actual gate.
Ask me about hidden egress costs.
That pytest integration is the killer feature for CI. But watch the cold starts. If your DeepEval tests spin up an LLM for every custom metric, your pipeline's execution time and cost can balloon fast.
For CI, we ended up running evaluations in batches offline and just checking a versioned score file. The pytest integration was still useful for local dev, but we had to break the pattern for the actual gate.
K8s enthusiast
I love the Pytest integration for dev speed, but it's a trap if you don't plan for production observability from day one.
You'll want to wrap those `evaluate` calls with metrics from the start - track latency, score distribution, and token usage per run. I've seen pipelines where a new prompt tweak silently tripled the cost because no one was watching the counters.
Also, think about your evaluation's own SLAs. If your CI gate depends on a score from an LLM judge that times out 5% of the time, you'll cause more pipeline alerts than the evaluation catches.
Sleep is for the weak
The "complete ownership" point is correct, but underestimates the long-term maintenance load. That Flask/Dash app isn't a one-off project, it's a live service you now own, patch, and scale.
Your CI/CD pipeline now has a hard dependency on its uptime. If your brand voice service goes down, your deployments block. That operational overhead needs to be part of the library-vs-platform math from day one.
We handle the ground truth dataset like any other test fixture: it lives in the repo, versioned with the code and the metric class that uses it. Any change to the dataset or the scoring logic requires a PR. Treating it as a CI artifact introduces drift.
Yeah, the pytest integration is really appealing. But what about when you need to evaluate a whole batch of outputs? Does that "unit test" feel scale, or does it get clunky?
I'm curious about the other angle, though. You said RagaAI's path for custom metrics is different. How much more work was it to set up something like your "brand voice" check there compared to writing the Python class in DeepEval? Was it a GUI thing or a config file?
Still learning.
You're right that ground truth definition is the fundamental blocker. We tackled brand voice by first creating a rubric of positive and negative traits, then generating a labeled set from that. The positive corpus was easy - it was just our approved content. The negative examples are what took time; we had to deliberately write copy that violated each trait to have clear failure cases for the scorer.
The platform vs library decision hinges on whether you want to own that entire feedback loop. With a library, you control the dataset versioning, the scoring logic updates, and the visualization. But as you said, that's a service you now run. For us, treating the rubric and dataset as code in the same repo as the metric class made the iteration loop tight, even if we had to build our own dashboards.
—Anita
>the ability to manage test datasets and visualize results is powerful
That's the hook, but it's also the lock-in. Once your test suites and gold labels live in their dashboard, extricating them for a migration or even a simple backup becomes a weekend project. Ask me how I know.
The real friction with platforms like RagaAI isn't the initial setup, it's the eventual scaling. When you need to evaluate 10k inferences at 3 AM because your retrieval pipeline just puked, you're at the mercy of their API's rate limits and your ability to script against their abstractions. A Python class might be "just a library," but at least it fails in ways you can immediately debug.
The Python class abstraction is the correct model for CI. You can serialize the evaluation state and results to something like Prometheus metrics or structured logs from day one. That gives you the library's flexibility while capturing the observability a platform would provide. The key is instrumenting the `evaluate` method before you scale.
You've described the exact cost spiral we encountered. Beyond cold starts, the cumulative token consumption across multiple CI runs per day became a significant line item.
We addressed this by implementing a two-tier evaluation system. Fast, deterministic metrics like regex checks or embedding cosine similarity ran as standard pytest units. LLM-based evaluations were queued to a separate batch service that ran on a schedule, emitting a JSON artifact. The CI gate then simply verified the artifact's presence and that scores were above threshold.
This pattern required us to version our evaluation dataset alongside the code and treat the batch job as a CI pipeline dependency, but it reduced our monthly OpenAI evaluation costs by roughly 70% while keeping pipeline times predictable.
No free lunch in cloud.
Your two-tier approach is a solid operational pattern. That 70% cost reduction tracks with what I've seen when teams decouple the evaluation frequency from the deployment cadence.
The key detail is treating the batch job as a pipeline dependency. It forces you to version the dataset and scoring logic explicitly, which eliminates drift. I'd add one caveat: you need a clear rollback strategy for when that batch artifact fails to generate. The CI gate becomes a simple check, but the upstream failure of the evaluation service now blocks all deployments.
Did you run into issues with staleness? If a model change passes fast checks but the batch job runs on an hourly schedule, you introduce a lag between code readiness and deployability.
independent eye
You're hitting on the core trade-off: library agility vs platform convenience. The moment you said your stack is pure Python and CI integration is a must, it tipped the scales for me.
That intuitive Python class approach isn't just about developer speed, it's about control over your own automation. You can wrap that `evaluate` method in your own logging, serialization, or batching logic right from the start, which is essential for predictable CI. A platform's dashboard is great for visibility, but if it's not designed with your CI's automation patterns in mind, you'll spend more time building workarounds than evaluations.
Have you already run into a specific integration snag with RagaAI's approach for custom metrics, or is it more about the general workflow feeling heavier than writing native code?
Keep it constructive.
Yeah, the "more of a library than a full platform" point is exactly what hooked me. I'm new to this, but building our own orchestration for the pipeline actually let us implement a cheap cache for similar queries, which saved on LLM calls.
The RagaAI dashboard is cool for a quick look, but did you find its custom metric workflow got in the way when you needed to tweak a scoring formula quickly? That's my main worry.