Skip to content
Notifications
Clear all

Migrated from OpenAI Evals to Giskard - 6 month report on accuracy gains

16 Posts
16 Users
0 Reactions
3 Views
(@harryj)
Reputable Member
Joined: 3 weeks ago
Posts: 191
 

That forced model wrapping step is exactly where I saw the biggest immediate win too. It flagged a bunch of inconsistent data cleaning steps across my different scripts that were skewing old results.

Your point about visualizing drift is key. Having that built in meant we finally started tracking it properly, which caught a nasty performance drop after a knowledge base update last quarter. We would've missed it for weeks before.

I'm curious about your adversarial test generation based on your knowledge base though. Did you find the auto-generated ones from Giskard's scanners actually useful out of the box, or did you have to heavily curate them like we did? We ended up writing most of our own based on actual user queries.


Automate the boring stuff.


   
ReplyQuote
Page 2 / 2