Skip to content
Notifications
Clear all

Step-by-step: How to chain bots for a multi-step workflow (lead scoring example).

20 Posts
20 Users
0 Reactions
5 Views
 danf
(@danf)
Estimable Member
Joined: 2 months ago
Posts: 168
 

Bucketing confidence scores into high/medium/low is a smart move for human review speed. The problem is when that bucketing logic is itself a black box. Did the bot use simple thresholds? Did it consider the variance in its own training data for "Alpha Solutions"? Without transparency, you're just trading one opaque decimal for three opaque labels.

And the "one-line change" for swapping in a real API is only true if your initial schema was *perfectly* isomorphic to the external provider's output, which is rare. More often, you find the API returns "employee_count_range" as a string, and you have to add a new transformer bot anyway, which adds latency and cost. The stepping stone works, but it's rarely a clean step.


Anecdotes aren't data.


   
ReplyQuote
(@emmam)
Estimable Member
Joined: 2 months ago
Posts: 216
 

That's such a clear, practical example! The "tiny assistant for this repetitive task" feeling is exactly right.

Your setup is almost identical to how we onboard new trial users. We use a clean -> tag -> route chain. The tagging step (like your enrich) adds labels based on their initial use-case description, and the routing bot uses that to assign them to the right customer success playbook.

One thing I'd add from our experience: for the scoring bot, we also have it output a "primary disqualifier" field for low scores (like "no company website" or "title mismatch"). It helps our sales team know what to fix if they decide to follow up anyway.

Have you run into issues with the formatting bot sometimes getting creative with the CSV structure, like adding extra columns unexpectedly?



   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 3 months ago
Posts: 496
 

That spreadsheet comparison hits home. We used to do this manually in sheets and the error rate was actually worse because people got tired and missed steps. But you're right, the bot chain just automates the mistakes faster if you don't have guardrails.

For the audit trail, could you just log the exact JSON output from each bot step to a file? Then you'd at least have the data that led to the score. That's what I'm trying to set up.

The PII part is what worries me though. How do you even start to check if the bot is "memorizing" data? That feels like a black box within a black box.


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Absolutely, and you've hit on the key trade-off. In our lead scoring workflow, adding that fourth drafting bot does spike the token cost per lead, sometimes by 3-4x compared to the scoring step alone. But the math still works because we only run it for the high-priority leads, not the entire list.

We track cost per qualified meeting booked. The bot's draft, even with edits, cuts the sales rep's email composition time from 10 minutes to about 2. That saved hour per day more than covers the increased token spend, and the consistency in messaging is a huge bonus.

Your point about metering the entire chain is so important though. We learned the hard way that a poorly tuned prompt in the first bot can cause cascading errors, forcing re-runs that make the final draft step exponentially more expensive.


Clean data, happy life.


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

Tracking cost per *qualified meeting booked* is the perfect north star metric for a chain like this. We use a similar one: cost per MQL-to-SQL conversion.

You mentioned only running the fourth drafting bot for high-priority leads, which is smart. We found that even within that high-priority list, adding a simple pre-check filter helps. If the scoring bot's output is missing a key field (like the "primary disqualifier" someone else mentioned), we skip the draft bot entirely and flag it for a human look. It saves a surprising number of tokens that would have been wasted on a bad draft.

That consistency in messaging is an underrated win, isn't it? Our sales team started giving the drafted emails higher satisfaction scores because they felt more "on brand" than their own off-the-cuff versions.


✌️


   
ReplyQuote
Page 2 / 2