Skip to content
Notifications
Clear all

Guide: Benchmarking output quality for sales email generation across 5 providers.

2 Posts
2 Users
0 Reactions
1 Views
(@ethans)
Estimable Member
Joined: 2 weeks ago
Posts: 48
Topic starter   [#22377]

Just tried generating a sales email for a new analytics SaaS feature across 5 providers. Same system prompt and prospect details each time. The goal was a short, personalized cold email. Results were surprisingly varied.

Claude gave the most natural, consultative tone but was a bit long. GPT-4 was sharp but felt generic. GPT-3.5-Turbo was fast and cheap, but the quality drop was obvious—too pushy. Jurassic-2 was formal and off-brand for a startup. Command R+ had good structure but missed some personalization cues. For cost vs. quality balance, I'm leaning towards Claude for high-touch leads and GPT-3.5 for volume, but still testing. Anyone else run similar side-by-sides?



   
Quote
(@data_meets_ops)
Estimable Member
Joined: 2 months ago
Posts: 93
 

Interesting breakdown. The "consultative tone but a bit long" observation on Claude is spot on - I've found that adding a specific instruction for a maximum sentence count or word budget in the system prompt can sometimes reel it in without losing that quality.

For high-volume scenarios where you're considering GPT-3.5, have you tried running its output through a separate, simple checker for personalization tokens? We built a lightweight step that flags if company name/role/mentioned pain point are missing, which catches a lot of the generic feel before sending. It's a cheap gatekeeper.

Your split strategy makes sense. Are you tracking open/reply rates per provider to see if the tone difference actually translates to performance?



   
ReplyQuote