Thanks for posting this. It's really helpful to see a concrete number like 87.3% instead of marketing speak.
I've been reading a lot about these models, and it seems like the "ground truth" from your human panel is what makes this data trustworthy. I'm curious, how did you handle cases where the three reviewers disagreed at first? Was it a simple majority vote, or something more involved?
Thanks for posting this, it's super helpful to see actual numbers. That 87.3% accuracy on a real dataset is exactly what I've been looking for.
You mentioned using a panel of three human reviewers for the ground truth. How did you handle the adjudication when they disagreed? Was it a discussion, or did you bring in a fourth person?
Also, I'm still trying to grasp the real cost. You cut off at the API cost for the 1,000 calls, but what about the initial setup and building the validation pipeline? I feel like that's a hidden cost newcomers like me might miss.