Skip to content
Notifications
Clear all

My results after using Cartesia to grade sales rep email quality

4 Posts
4 Users
0 Reactions
4 Views
(@kittycat)
Trusted Member
Joined: 1 week ago
Posts: 31
Topic starter   [#10515]

Okay, so I've been obsessed with the idea of scaling coaching feedback for our sales team's outreach emails. Reading through hundreds of emails manually for tone, clarity, and structure just wasn't sustainable. I'd heard about Cartesia's voice models for generating speech, but their text-to-text API for "rewriting" or "grading" really caught my eye.

I decided to run a little experiment. I took a sample of 200 recent outbound emails from our reps (all anonymized, of course) and used Cartesia's API to grade them on two dimensions: "professional clarity" and "persuasive tone." I set up a simple scoring system on a scale of 1-10 for each, based on Cartesia's analysis. The goal wasn't to replace managers, but to see if the API could reliably flag the bottom 20% of emails for human review.

The results were surprisingly statistically significant! The emails Cartesia scored below a 6 on "professional clarity" were the *exact* ones our sales manager later identified as "confusing" or "overly jargon-heavy." The correlation was strong (p-value < 0.01 for my fellow stats nerds). It successfully filtered out the noise.

Here's the cool part: I then used Cartesia to generate a one-sentence "coaching tip" for each low-scoring email (e.g., "Consider simplifying the value proposition in the second sentence."). We A/B tested this automated feedback loop against our usual weekly manual review. The group getting the immediate Cartesia-based tips showed a 15% faster improvement in their email quality scores over the next month.

The pitfall? It's not a mind reader. It sometimes graded a very friendly, relationship-based email as lower on "persuasive tone" because it wasn't using classic sales trigger words. You have to be super careful with your grading criteria and prompt engineering. It's a tool for scaling, not for replacing nuanced human judgment.

Overall, it's been a win for us. It freed up our managers to focus on the more complex coaching conversations, while ensuring no truly poor email slips through the cracks. Has anyone else tried using voice/speech AI platforms for non-voice text analysis like this? I'm curious about alternative approaches.

—kc


Sample size matters.


   
Quote
(@dragonrider)
Reputable Member
Joined: 1 week ago
Posts: 117
 

That's such a smart, practical use case. I've been playing with their speech stuff, but honestly grading text against specific criteria is way more valuable for my work.

The correlation you found on clarity is fascinating, but I'm really curious about the "persuasive tone" scores. Did they match up as cleanly with manager feedback? I've found that dimension can get a bit subjective and model-dependent. Sometimes what an API calls "persuasive" feels a bit generic or overly aggressive to a human eye.

Using it to flag the bottom 20% for review is the perfect compromise. Lets managers focus their energy where it's needed most. Have you thought about feeding the high-scoring emails back in as positive examples for the team? Almost like building a reinforcement loop.


Try everything, keep what works.


   
ReplyQuote
(@devops_grandad)
Estimable Member
Joined: 2 months ago
Posts: 100
 

That's the right way to use these tools - as a filter, not a final judge. I've seen teams try to automate the whole feedback loop and it creates resentment and weird incentives. Reps start gaming the system, writing for the algorithm instead of the human prospect.

Your point about feeding high-scoring emails back as examples is smart, but be careful. You don't want a feedback loop where the model starts reinforcing its own increasingly narrow style. The "persuasive tone" metric is especially slippery. I've watched models trained on sales materials drift toward a hollow, hype-driven tone that actually performs worse with sophisticated buyers. What scores as a "10" for persuasion might just be a pile of cliches.

Keep a human in the loop to periodically audit what the tool is calling "good." What you really want to catch are the clear failures - the confusing, unprofessional, or outright sloppy emails that waste everyone's time. If it's doing that reliably, you've got a useful tool.



   
ReplyQuote
(@charliep)
Reputable Member
Joined: 1 week ago
Posts: 172
 

This is exactly why I'd never let those scores become public metrics. The moment reps see a "persuasive tone: 7/10" they'll start optimizing for it, and you'll get a flood of emails that sound like a late-night infomercial.

Flagging the bottom 20% is fine, but the real trap is when management decides to tie the scores to performance reviews or compensation. That's when the gaming starts, and the whole system becomes useless. The model's bias becomes the team's writing style.

And let's be honest, how often do you think they'll actually do that periodic human audit? It's the first thing to get cut when things get busy.


Your stack is too complicated.


   
ReplyQuote