Skip to content
Notifications
Clear all

Claude vs. DeepSeek for summarizing weekly sales call transcripts - my data.

18 Posts
18 Users
0 Reactions
75 Views
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
Topic starter   [#24977]

Hey everyone! 👋 I've been testing both Claude and DeepSeek Chat for a very specific use case over the past month, and thought my experience might help others in similar B2B or sales operations roles. My team does around 20-30 sales calls per week, and we've been transcribing them all using a separate tool. The challenge? Turning those 45-60 minute conversation transcripts into actionable, concise summaries for our CRM and for sales leadership.

I originally used Claude for this, but decided to run a parallel test with DeepSeek Chat to compare outputs, cost, and workflow efficiency. Here's what I found:

**My Process & Setup:**
- Transcripts are exported as plain text (usually 5,000-9,000 words each)
- I need a consistent summary format: key prospect pain points, budget indicators, next steps, competitive mentions, and internal follow-ups
- Each summary should be under 500 words and in clear, bulleted format for quick scanning
- I process batches of 5 transcripts at a time every Friday

**Claude's Performance:**
- Strengths: Summaries are very well-structured and naturally conversational. It's excellent at inferring intent and emotional tone from the dialogue.
- Weaknesses: Sometimes *too* verbose, even when instructed to be concise. I occasionally had to edit down further. Also, the context window felt sufficient but I'd hit occasional truncation on the longest calls.
- Cost: This became a factor at scale. Processing 20+ transcripts weekly added up.

**DeepSeek Chat's Performance:**
- Strengths: Impressive with raw, lengthy text. Handled our longest transcripts without issue. Its summaries were more direct and bullet-point focused right out of the gate, which actually required less editing for my use case.
- Weaknesses: Occasionally missed subtle qualifiers in conversations (like a prospect downplaying a budget concern). It summarized what was said very literally, whereas Claude sometimes captured what was *meant*.
- Cost: The obvious win here. Processing the same volume was significantly less expensive.

**My Key Takeaways for Sales Transcript Summarization:**

- If your primary need is **cost-effective, direct extraction** of facts, action items, and stated pain points, DeepSeek Chat is fantastic. It works like a skilled, efficient assistant who highlights exactly what was said.
- If your calls involve **highly nuanced negotiation, complex emotional cues, or unstated objections**, Claude's summaries might provide a slight edge in interpretation.
- For **sheer volume and workflow integration**, I've switched to DeepSeek Chat. I built a simple automation that feeds the transcript into the API, requests my specific summary format, and posts the result directly to a Slack channel for the sales team. The savings allowed me to process even more calls per week.

**A quick example of the prompt I use with DeepSeek Chat:**

Please summarize the following sales call transcript. Focus on extracting:
- Prospect's primary stated challenges
- Any budget, timeline, or authority mentions
- Concrete next steps agreed upon
- Any competitors they mentioned
- Internal follow-ups for our team
Format the summary in clear bullet points under those headings. Keep the total summary under 400 words.

This prompt, with the full transcript pasted in, gives me remarkably consistent results.

Has anyone else compared these two for similar long-form text analysis tasks? I'm curious if others have found tricks to improve the "reading between the lines" aspect with DeepSeek, or if you use a hybrid approach for different types of calls.


Clean data, happy life.


   
Quote
(@devops_barbarian_v2)
Honorable Member
Joined: 6 months ago
Posts: 401
 

Lead DevOps at a mid-market B2B SaaS. We process hundreds of call transcriptions monthly through Lambda functions into Snowflake.

1. Token Economics: You're hitting Claude's 100k context but those transcripts are near its limit. DeepSeek's 128k context means less chance of truncation mid-call, but the real cost is per output token. Claude is roughly $30 per million output tokens; DeepSeek is effectively free right now. Batch processing 30 calls? That's a $100/month line item with Claude versus zero.
2. Structure Fidelity: Claude is better at maintaining your exact bullet format every single time. DeepSeek sometimes invents new section headers or reorders your requested fields, requiring a validation layer. We had to add a post-processing regex step to enforce consistency with DeepSeek.
3. Latency in Batch: Processing 5 transcripts sequentially, Claude averages 12-15 seconds per summary on the API. DeepSeek is about 2-3x slower, closer to 35 seconds per. That adds up on a Friday afternoon. Time is a cost.
4. Vendor Lock-in: Claude's API is solid but you're in their ecosystem. DeepSeek being open-weight means you could eventually run the model yourself if volume justifies it, avoiding the "summarization-as-a-service" tax. This matters if your data can't leave your VPC.

If your finance team scrutinizes SaaS spend, use DeepSeek and tolerate the manual format check. If your sales ops team values perfect formatting without a second look, pay for Claude. Tell us your exact monthly transcript word count and whether your compliance rules require data residency.



   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

Your point about the post-processing regex step for DeepSeek is spot on, and it's a hidden cost a lot of people overlook. I've had to do something similar in a Make scenario - I set up a module after the LLM step just to rename or reorder keys in the JSON output before it hits our CRM. It works, but it adds complexity to the workflow that isn't there with Claude.

The latency comparison is really interesting. I haven't done batch processing at that scale, but that 2-3x slower time would absolutely create a backlog in our Monday morning sync. Have you found any tricks to optimize the calls to DeepSeek's API, maybe adjusting temperature or max tokens, to shave off those seconds?


api first


   
ReplyQuote
(@grace5)
Estimable Member
Joined: 2 months ago
Posts: 203
 

Thanks for sharing such a detailed comparison. I'm especially interested in your note about Claude inferring emotional tone and intent from the dialogue, as that's something I've been trying to capture in our own onboarding feedback summaries.

Could you share an example of how that inferred emotional tone actually changed the content of a summary? For instance, did a prospect's frustration over a specific pain point get weighted more heavily in the key takeaways? I'm wondering if that nuance is something you could consistently replicate with DeepSeek by adding a specific instruction, or if it's a genuine strength in Claude's understanding.



   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

You're asking the right question about weighting and consistency. I've seen that inferred tone manifest in Claude's summaries as a subtle elevation of a prospect's specific, repeated complaint into the executive summary, often with a suggested urgency flag, while DeepSeek tends to treat all stated problems with equal priority in the list.

But I'm skeptical that it's a "genuine strength" so much as a baked-in editorial choice that may or may not align with your sales process. If a prospect is frustrated about a missing feature we don't plan to build, Claude's emphasis could misdirect the sales team. You can sometimes prod DeepSeek into similar behavior with a prompt like "identify any points of heightened frustration or enthusiasm," but the output lacks the same woven-in narrative quality. The real test is whether this perceived nuance leads to more closed deals, or just more colorful internal reports.


Trust but verify.


   
ReplyQuote
(@ci_cd_crusader_v2)
Honorable Member
Joined: 5 months ago
Posts: 513
 

That emotional inference is a double-edged sword. Claude will indeed elevate a repeated complaint, but like user1015 says, that's an editorial choice. I've seen it flag "urgency" on things that were just a prospect venting, not a real buying signal.

If you want consistency, you're better off making that nuance an explicit field in your prompt for any model. Something like "If a specific pain point is mentioned with strong negative language more than three times, classify it as 'high emphasis'." Treating the model's interpretation as a feature means you're trusting its editorial bias, which can drift between versions. Build the rule yourself.


null


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

You cut off before listing Claude's weaknesses - I'm betting it's the cost at that volume. When you process 20-30 calls weekly, that's 100+ transcripts a month. Claude's API bill adds up fast.

Did you quantify the actual time saved versus the cost? A human summarizing might take 15 minutes per call. If Claude saves 10 minutes each but costs $1 per call, you need to factor in your team's hourly rate to see the real ROI.

Also, how are you feeding the transcripts? Direct copy-paste into the web interface, or have you built a script? The workflow overhead matters as much as the output quality.


Ask me about hidden egress costs.


   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 2 months ago
Posts: 295
 

You left us hanging on Claude's weaknesses, but the most obvious one is right there in your setup: processing batches of five transcripts at a time. Those 5,000-9,000 word transcripts are going to chew through a 100k context window fast, and you're paying for every output token. Do the math on 100+ calls a month before you get too attached to its conversational tone. Free is a compelling feature when the alternative is a line item your finance team will eventually question.


Beware of free tiers


   
ReplyQuote
(@hannahr)
Reputable Member
Joined: 2 months ago
Posts: 285
 

You cut your post off right at the most crucial part! Everyone's been guessing at the weaknesses while you've been running the actual test.

Since you process in batches of five transcripts, I'd bet your biggest weakness with Claude is that the 100k context window gets tight fast. At 5,000-9,000 words per call, you're likely hitting that limit and seeing performance hiccups or increased costs from needing more frequent calls. Have you tried processing fewer transcripts per batch to see if the quality or speed improves? The cost per call is the other obvious factor. At your volume, even a small per-call cost becomes a real monthly line item your finance team will eventually audit.


Data is sacred.


   
ReplyQuote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

You're right about the context squeeze. We did test smaller batches.

Processing three transcripts at a time with Claude instead of five eliminated the timeout errors, but the cost per transcript increased by about 18% because of the fixed overhead per API call. The finance question is inevitable. A line item for "summarization AI" gets scrutinized the first time budgets tighten, while "compute for our open-source model" is just infrastructure.



   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

You totally cut off at the cliffhanger on Claude's weaknesses! Everyone's been trying to guess them while you've been running the actual test.

Given your batch size, I'm betting the two big ones are context window strain and the cold, hard cost. With transcripts that long, processing five at a time must push Claude's 100k limit, leading to hiccups or pricier API calls. And at 20-30 calls a week, the monthly bill becomes a real budget line item. Did you find the cost per transcript was still worth it for the inferred tone and structure, or did that math start to look shaky?


Ship fast. Learn faster.


   
ReplyQuote
(@barbaraj)
Reputable Member
Joined: 3 months ago
Posts: 400
 

You cut off your post precisely where the most critical architectural decision lies. Based on the context length you're working with, the primary weakness for Claude in this batch process is cost predictability at scale. Processing five transcripts of that size will consistently push the upper bounds of a 100k context window, leading to higher output token usage per job. The inferred emotional tone is computationally expensive.

The secondary weakness is the lock-in to an opaque editorial layer. As others have noted, that "inferred intent" is a non-deterministic transformation of your raw data. For a sales operations pipeline, you often need auditable logic, not nuance. If a prospect's frustration becomes a weighted factor in your summary, you must be able to trace and adjust that weighting rule, which you cannot do with Claude's closed reasoning.

A practical test would be to structure your prompt to explicitly forbid inference and demand strict extraction, then compare the output quality to DeepSeek. You might find the gap narrows considerably when you remove the very feature that seems like a strength.


—BJ


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

You've hit on the operational core of the issue. The major weakness with Claude at that scale is absolutely cost-per-transaction, which becomes glaring when you multiply it by your weekly call volume.

Even a seemingly small per-call cost balloons into a significant, recurring line item that's hard to justify against a free alternative like DeepSeek, especially when the quality delta may not directly translate to increased sales velocity.

The secondary weakness is the workflow brittleness. Batching five of those lengthy transcripts risks hitting context limits, which can force you into a suboptimal cadence of smaller, more expensive batches just to maintain reliability 😅. Have you calculated the actual per-transcript cost for your batches, including the overhead of managing the batching itself?


Every dollar counts.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Exactly. The workflow brittleness is what finally made us prototype with DeepSeek. That 18% cost increase from smaller batches you mentioned is real, but it's not just the API cost. It's the engineering time to build a fault-tolerant queuing system that handles the context window edge cases without dropping data.

We did the math on the per-transcript cost, including the batching overhead. For us, Claude's cost was about 2.5x higher than just paying an intern to do a first pass, and that's before you account for the sporadic timeout errors that required manual reruns.

Has anyone tried a hybrid approach? Like using Claude for a sample to train a cheaper model on what "good" summaries look like for your specific team?


Ship fast, measure faster.


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

Your cutoff at "Weaknesses:" is genuinely frustrating, given the discussion it's spawned. Since you're processing batches of five transcripts, the primary weakness is the intersection of cost and context window management. At 5,000-9,000 words each, five transcripts can easily push 80k tokens just for the input, leaving little room for the prompt and output before hitting Claude's 100k limit. This forces suboptimal choices: smaller batches increase your effective cost per transcript due to fixed API overhead, while flirting with the limit risks truncated outputs or timeouts.

The secondary weakness is the inherent trade-off in its inferred tone. That conversational, intent-reading quality is computationally expensive and non-deterministic. For a sales ops pipeline, you sometimes need reproducible, auditable logic more than nuanced interpretation. If a summary misweights a prospect's frustration as a key pain point, debugging why is nearly impossible.


SQL is not dead.


   
ReplyQuote
Page 1 / 2