Skip to content
Notifications
Clear all

Check out this comparison table I made for output on sales email templates.

2 Posts
2 Users
0 Reactions
1 Views
(@jakem)
Estimable Member
Joined: 1 week ago
Posts: 72
Topic starter   [#13874]

I've been evaluating LLM assistants for sales team use cases, specifically for generating and refining outbound email templates. The key requirement is consistent, professional output that adheres to brand voice—without excessive creativity that derails the message.

I ran a structured test comparing three models (Claude 3 Opus, GPT-4 Turbo, and DeepSeek-V3) on the same prompt: "Generate a concise, value-based outreach email for a cloud cost management platform aimed at a Director of Engineering." I then scored the outputs across five criteria critical for sales ops.

| Criterion | Claude 3 Opus | GPT-4 Turbo | DeepSeek-V3 | Notes |
|-------------------------|---------------|-------------|-------------|-------|
| Adherence to Prompt | 9/10 | 8/10 | 7/10 | DeepSeek added a slightly more technical intro than requested. |
| Brevity & Conciseness | 8/10 | 7/10 | 9/10 | DeepSeek's output was notably tighter, under 120 words. |
| Perceived Professional Tone | 9/10 | 8/10 | 8/10 | All were acceptable; Claude edged out on formal phrasing. |
| Action-Oriented CTA | 8/10 | 9/10 | 7/10 | GPT-4 had the strongest call-to-action structure. |
| **Cost per 1k Output Tokens** | **$0.075** | **$0.03** | **~$0.0007** | Estimated based on current pricing; DeepSeek is orders of magnitude cheaper. |

**Key Takeaway:** For high-volume, templated content generation where absolute perfection isn't required, DeepSeek presents a compelling TCO argument. The quality delta isn't proportional to the cost difference. For a sales team generating hundreds of email variations weekly, the cost savings could be redirected to more human refinement or other tools.

However, for mission-critical, client-facing communications where nuance is paramount, the more expensive models might still justify their premium. It's a classic build vs. buy (or rather, cheap vs. expensive) calculation.

Has anyone else done similar A/B testing for operational content? I'm particularly interested in how consistent DeepSeek is over thousands of generations.

—Jake


Show me the bill.


   
Quote
(@devops_shift_worker)
Estimable Member
Joined: 2 months ago
Posts: 104
 

Nice comparison table, but you're missing a key column: prompt engineering overhead.

Claude might score a 9/10 on "Adherence to Prompt," but in my experience, getting *truly* consistent output from any of these models requires you to bake your brand voice and template rules into a system prompt that's longer than the email itself. Then you have to maintain that prompt across updates.

DeepSeek being tighter is a valid win, though. Our sales dev team complains when drafts run long. Maybe the trick is to start with the concise model and just polish the tone, rather than starting with the "most professional" output and trimming it back.


NightOps


   
ReplyQuote