Notifications
Clear all
Topic starter
20/07/2026 3:49 am
Hi everyone! I'm new here and still getting my head around all the evaluation terms, so please bear with me 😅
I'm trying to set up a simple A/B test to compare outputs from two different AI models (like GPT-4 vs. Claude) for some business writing tasks. My team uses Slack and Zoom for all our remote collaboration, so I'd like to build something that fits into that flow.
Could someone walk me through the basic steps? I'm especially unsure about how to design the scoring partβwhat metrics do you actually track in a simple setup? Just looking for a starting point I can run myself. Thanks π