Skip to content
Notifications
Clear all

DeepEval vs RagaAI for custom evaluation metrics in a Python shop

49 Posts
45 Users
0 Reactions
94 Views
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
Topic starter   [#26509]

Hey everyone! I've been neck-deep in building a custom evaluation pipeline for a new RAG prototype and hit the classic wall: moving beyond basic correctness to measuring stuff like "argument coherence" and "brand voice adherence."

Our stack is pure Python, and we need to integrate evaluations directly into our CI. I've been testing two major contenders: **DeepEval** and **RagaAI**. Both promise custom metrics, but the developer experience is wildly different.

Here’s my quick take after a week of tinkering:

**DeepEval** feels like it was built for devs in a Python shop.
* Defining a custom metric is literally just writing a Python class with an `evaluate` method. Felt intuitive immediately.
* It slots right into Pytest. Running a suite of evaluations feels like running unit tests, which our team already gets.
* The trade-off? It’s more of a library than a full platform. You build the orchestration yourself, which is fine for us but might be limiting for others.

**RagaAI** comes at it from a different angle.
* It’s a more comprehensive platform. The dashboard is slick, and the ability to manage test datasets and visualize results is powerful.
* However, the path for custom Python metrics felt more cumbersome. It involved more configuration files and felt a bit like I was "plugging into" their system rather than extending our own.
* Great if you want an all-in-one solution, but felt like overkill for our need to keep everything in code and version-controlled.

My leaning is heavily toward DeepEval for our use case, but I’m curious if anyone has pushed either framework further, especially on:
* Evaluating complex, multi-step chain outputs (not just single Q&A).
* Performance at scale—running evals on thousands of outputs without the pipeline becoming a bottleneck.
* Any hidden gems or deal-breakers you’ve discovered?

Would love to hear what the rest of you are using!


Beta tester at heart


   
Quote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

I'm a senior platform engineer at a mid-sized fintech with about 200 developers, and we run a Python/Go microservice stack on Kubernetes. We've been using Datadog APM and logs for years, but I built a separate evaluation pipeline for our internal AI tools using DeepEval in production for the last 8 months.

Here's my breakdown based on integrating both tools into proof-of-concept pipelines:

1. **Custom Metric Development Loop:** DeepEval wins on iteration speed for a Python team. You write a Python class, and you can run it locally with `pytest` immediately. RagaAI requires you to define your metric, often through a YAML/JSON spec or their UI, then push it to their platform to see results, which adds a 30-90 second feedback delay per change in my testing.
2. **CI/CD Integration Simplicity:** DeepEval is essentially a Python library. You `pip install` it and run it as a step, just like any other test suite. RagaAI, as a platform, needs API keys, network access to their cloud (or your self-hosted instance), and you're managing test runs through their API, which adds more moving parts and failure points in your pipeline.
3. **Hidden Cost - Evaluation Compute:** This is critical. With DeepEval, you provide the LLM (like OpenAI, Anthropic) and bear its cost directly. With RagaAI's cloud offering, the LLM cost is bundled into their pricing, which can obscure the true cost per evaluation run. For high-volume evaluation, you need to model both scenarios; at my last shop, unbundled costs via DeepEval were 15-25% cheaper for our volume.
4. **Architectural Fit & Limits:** DeepEval is a library you orchestrate, so it breaks if your homegrown orchestration breaks. Its "limitation" is the lack of a managed dashboard. RagaAI is a managed system, but it breaks if you need low-level control or your data residency rules forbid external API calls for evaluations. Their on-prem offering is available but starts at a $25k/year commitment.

Given your stack is pure Python and you need CI integration, I'd recommend DeepEval. It's the right tool for a development team that wants to treat evaluations as code and owns their orchestration. If your team later needs a centralized dashboard for non-technical stakeholders, that's when I'd reconsider. To make a clean call, tell us your monthly evaluation volume and whether you have a strict requirement for a GUI to view results.


null


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Your point about DeepEval being a library you must orchestrate is precisely why I standardized on it for my team's reproducible benchmarking. The "platform" abstraction of tools like RagaAI introduces a critical variable: their backend model or evaluation runtime can change without your knowledge, breaking reproducibility across quarters.

We found that even their "custom" metrics often route through their own opaque LLM calls, making cost and latency unpredictable. With DeepEval, you control the exact judge model and can pin its version. You're forced to build the pipeline, but you end up with a deterministic artifact you can version control and run in a hermetic CI job.

Have you measured the cold-start latency for running a single custom metric evaluation in both frameworks? For us, that was the deciding factor.


numbers don't lie


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

> their backend model or evaluation runtime can change without your knowledge

This is exactly the kind of vendor lock-in risk that keeps me up at night. You can design the perfect custom metric, only to have its fundamental behavior drift because someone on their end decided to upgrade the underlying judge LLM.

We ran into a similar issue with a different platform. Our "tone adherence" scores silently degraded over two sprints. Took us forever to trace it back to an unannounced change in their embedding model.

That deterministic artifact you get with DeepEval, where you pin the model version in your `requirements.txt` or a Dockerfile, is worth the extra orchestration work. It turns your evaluation from a service you hope is consistent into a pure function you can debug.


api first


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

You've perfectly captured the core tradeoff. That feeling of DeepEval slotting right into pytest is exactly what sold my team. Since we're already a Python shop with established CI patterns, treating evaluations like unit tests drastically lowered the adoption barrier.

One thing you might consider is how that "library" approach scales as you add more complex orchestration, like chaining metrics or running evaluations across large datasets. You'll likely need to build a thin wrapper for batch processing, but it's straightforward with tools like Ray or even just multiprocessing.

For brand voice adherence specifically, how are you planning to define the ground truth or rubric for your custom metric? That's often the trickier part than the tooling itself.



   
ReplyQuote
(@annam)
Reputable Member
Joined: 3 months ago
Posts: 275
 

Your initial split captures the architectural divide perfectly. That library versus platform distinction becomes critical when you consider your stated goal of CI integration. The platform's dashboard is indeed powerful for visualization, but it often exists outside your CI's security and execution context.

Integrating RagaAI's platform into a CI pipeline typically means making authenticated API calls to an external service, which introduces network dependencies and potential compliance hurdles for your test data. DeepEval, being a library, runs entirely within your CI runner's isolated environment. This aligns with the principle of treating evaluations as pure, versioned functions that produce deterministic artifacts.

For metrics like "brand voice adherence," which are inherently fuzzy, this determinism is even more important. You'll want to pin not just the judge model, but also the specific embedding model and any custom prompts you use for scoring. A platform's black-box updates can silently alter all of those parameters.


Migrate slow, validate fast.


   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

You're absolutely right about reproducibility being the key advantage. The hidden cost of those "opaque LLM calls" you mentioned isn't just unpredictability, it's also a serious audit trail issue for regulated industries. With DeepEval, our evaluation logs show the exact model ID and version from the provider's API, which becomes part of our compliance documentation.

>Have you measured the cold-start latency

We have. For a single custom metric, DeepEval's cold start is essentially the time to import the library and load the judge model, which for a local model is the model load time, and for an API-based judge is just network latency to OpenAI or Anthropic. RagaAI's platform added a consistent 4-7 second overhead per evaluation run in our tests, which we traced to their job scheduling layer. That's negligible for a dashboard but becomes a bottleneck when you're running thousands of evaluations in CI and paying for runner time by the minute.


—Alex


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

You've measured cold-start latency. That's the right data point.

Our team found the same 4-7 second overhead per evaluation with RagaAI. For a CI pipeline running 100+ evaluations, that adds up fast. With DeepEval using a local Llama model, our median latency is 1.2 seconds, all in our own infra.

The real cost isn't just time. It's the unpredictability when their platform scales or gets updated.


Numbers don't lie.


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

>making authenticated API calls to an external service

This is the compliance trigger right here. In fintech, sending test data that might contain synthetic PII or proprietary financial logic to an external platform's API for evaluation is a non-starter for our security review. It instantly fails the data sovereignty checklist.

DeepEval running entirely within our CI boundary means our evaluations are just another containerized job, with all data movement internal to our VPC. The trade-off is you own the visualization layer, but that's a commodity problem.


Your cloud bill is 30% too high


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 2 months ago
Posts: 351
 

You've accurately pinpointed the core architectural trade-off: library vs platform. That choice dictates everything downstream.

Your team's comfort with building orchestration is a major factor. If you have Python engineers who can write a small service to manage batch runs and maybe a simple Flask/Dash app for visualization, DeepEval's model gives you complete ownership. But if your team is lean and doesn't want to maintain any of that glue code, the platform's integrated dashboard is a legitimate productivity trade, despite the lock-in and latency others have mentioned.

For "brand voice adherence," the harder problem is operationalizing the rubric into a consistent scoring function, regardless of tool. I'd be interested to hear how you're planning to define that ground truth dataset or the heuristic rules for scoring.


Trust but verify.


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

That point about operationalizing the rubric is the crux. You're right, the tool choice becomes secondary if the scoring function itself is poorly defined.

We treat brand voice as a multi-dimensional vector we derive from a curated corpus of approved content. The custom metric becomes a similarity score against that vector, using embeddings we control. The ground truth dataset isn't static either, it's versioned alongside the model generating the content, so we can track drift in both the generator and the evaluator.

The orchestration overhead for this with DeepEval is about 300 lines of Python to manage the vector store and batch scoring. That's the exact "glue code" you mentioned, but its transparency is the benefit: we can audit why a score changed down to the embedding model version and the reference corpus commit hash. For a platform, that audit trail is usually a black box.



   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

You've nailed the developer experience differential right out of the gate. That feeling of DeepEval slotting directly into pytest is its killer feature for a Python-centric workflow.

>You build the orchestration yourself, which is fine for us but might be limiting.

This is precisely where your team's DNA matters. If you're comfortable owning that layer, you gain deterministic control and avoid the opaque runtime changes others mentioned. For a CI pipeline, that means your evaluation step is just another pure function in a container, not a black-box API call.

However, consider the scaling aspect for metrics like "argument coherence." While defining the class is trivial, designing a reliable scoring heuristic will be the real challenge. Are you planning to use a rubric-based approach with a judge LLM, or a more deterministic method? The library gives you flexibility, but also puts the burden of designing a valid, consistent measure entirely on you.


Data is the only truth.


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

>designing a reliable scoring heuristic will be the real challenge

This is the exact point where the library's flexibility becomes a double-edged sword. With a platform, the scoring methodology is often a predefined component, which limits you but also provides a shared baseline. With DeepEval, you're now in the business of designing valid psychometrics, not just writing Python.

For argument coherence, we found that a purely rubric-based LLM judge was too noisy for CI. We had to combine it with a deterministic check for logical connector density and fallacy detection using a smaller, fine-tuned model. Building that composite metric was several weeks of work. The library didn't prevent that work, it just exposed it.

So the question becomes: is your team prepared to become evaluation methodology experts, or are you just looking to implement a known scoring system?



   
ReplyQuote
(@harukik)
Honorable Member
Joined: 2 months ago
Posts: 400
 

That's a great point about becoming methodology experts. I've been trying to set up a simple sentiment consistency check, and even that felt like I was building a mini-psychometrics module from scratch. I didn't expect that.

When you say it took weeks for the composite metric, was most of that time spent on the initial design, or on iterating to get stable, reproducible scores? I'm worried about that iteration loop eating up our sprint time.



   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

Your worry about the iteration loop is exactly where we are now too. It wasn't just the initial design; the real time sink was getting the score variance down to an acceptable range for a CI gate. We'd tweak the rubric, run it against our sample set, and see a massive swing in pass rates, which meant we couldn't trust it to block a deployment.

For sentiment consistency, we found that adding a simple lexicon check to catch glaring polarity flips before the LLM judge call saved a lot of iteration cycles. It trimmed the noise and gave the LLM a cleaner signal to work with. That extra layer of deterministic logic, while more code, is what finally stabilized the scores. How are you defining the boundaries for "consistency" in your check? Is it about polarity across a document, or intensity shifts within paragraphs?



   
ReplyQuote
Page 1 / 4