Skip to content
Notifications
Clear all

My bot for summarizing sales calls works great, but it's burning through my message caps. Alternatives?

4 Posts
4 Users
0 Reactions
2 Views
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
Topic starter   [#29287]

So I built this bot that listens to my sales calls, transcribes them, and spits out a summary with next steps. It’s actually shockingly good—catches action items the team missed, flags competitor mentions, the whole deal. It’s become a non-negotiable part of our workflow.

Here’s the rub: we’re a small team, but we do a lot of calls. The bot uses a hefty LLM call per summary, and we’re hitting Poe’s message caps almost weekly. The math is simple and painful: more successful calls = more bot usage = hitting the paywall faster. Feels like we’re being penalized for actually using the tool effectively.

I’m now staring down the Poe subscription upgrade path, which got me thinking: is this the most efficient way? The bot itself is just a prompt and some basic logic. The real cost is the LLM API call (GPT-4) and the audio transcription. Poe bundles it all up neatly, but the pricing is opaque and tied to messages, not tokens or minutes.

So I’m exploring the wilderness of alternatives. I could:
* Rebuild it directly on OpenAI’s API with Whisper for transcription, paying per use.
* Try a dedicated call-summary SaaS (but they never seem as flexible).
* Stick with Poe but try to optimize the prompt to be cheaper (less output tokens?).

Anyone else running into this scaling problem with Poe? Specifically for bots that do real “work” and aren’t just chatting? I love the platform, but the message cap model feels like it’s built for casual conversation, not actual utility bots that chew through context.


Data over dogma.


   
Quote
(@deborahw)
Reputable Member
Joined: 3 months ago
Posts: 358
 

The most efficient way? It's definitely not Poe once you hit any real volume. You're paying a massive premium for the convenience of their bundling, and their opaque "message" metric is designed to confuse unit economics.

Rebuilding on direct APIs is the obvious next step, but let me play devil's advocate. That Whisper+GPT-4 pipeline will have its own gotchas: Whisper's per-minute cost, managing two separate API calls, prompt engineering for consistency. But at least you'll see the actual token burn and can start optimizing there.

Before you jump, have you actually tried a brutally simple first pass with a cheaper model? Like, use GPT-3.5-Turbo to extract just the action items and competitor mentions from the transcript? Might cut your core LLM cost by 90% for 80% of the value. The "hefty LLM call" might be overkill for every single summary.


—DW


   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Poe's bundling is a tax on abstraction, and you've outgrown it. The moment you said "non-negotiable part of our workflow," you graduated to needing real cost per unit.

>the real cost is the LLM API

Exactly. So start measuring it directly. A direct Whisper+GPT-4 pipeline will instantly show you the token burn per call. You can then start cutting: cheaper model for first pass, truncating transcripts, caching repetitive competitor mentions.

The dedicated SaaS tools are garbage for flexibility. You'll end up hacking around their limitations within a month.

Just build the damn pipeline. It's a weekend project, and you'll know your actual burn rate by Monday. Then the optimization fun begins.


Prove it.


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Totally feel you on hitting the paywall faster because the tool is working! That's the worst kind of sticker shock.

I think your instinct to look at the direct APIs is spot on. When I was in a similar spot, switching from a bundled service to direct OpenAI/Whisper cut my cost by over half immediately, just because I could see the actual token counts and optimize. You mentioned your bot is "just a prompt and some basic logic" - honestly, that's perfect. Porting that over is maybe an afternoon.

One thing I'd add to your list of options: before you rebuild the whole pipeline, try a quick experiment. Feed your existing transcripts (just the text) through a cheaper model like GPT-3.5-turbo with the same prompt. See if the output is still usable for, say, 80% of calls. If it is, you could implement a simple rule: use the cheaper model for standard calls, and only trigger the "hefty" GPT-4 call for your most important deals or complex conversations. That alone might keep you under Poe's caps while you figure out the long-term move.


Backup first.


   
ReplyQuote