I've been using Poe to automate a lot of our small marketing team's manual work, and one of the most useful patterns I've set up is chaining bots for a multi-step process. It saves so much time. I'll walk through my real example for scoring new leads from a webinar.
The goal: Take raw attendee info, enrich it, then score the lead's fit automatically.
Here's my chain:
1. **Format & Clean Bot:** First, I paste the messy attendee list. A simple bot with instructions to reformat names, companies, and titles into a clean CSV.
2. **Enrich Bot:** I feed that output into a second bot. It's prompted to research the company domains (using its knowledge) and add fields like "Estimated Company Size" and "Industry."
3. **Scoring Bot:** Finally, the enriched list goes to a third bot. I gave it our ideal customer profile and scoring rubric. It outputs a final column with a score (1-10) and a short reason.
The key is using the "Continue in another bot" feature for each step. You can set it up as a workspace and run the whole sequence with one paste of the initial data.
It feels like having a tiny assistant that just handles this repetitive task. Anyone else building chains for their workflows? Would love to swap ideas.
Docs save time
Your approach is solid for a quick-start workflow. The manual handoff using "Continue in another bot" is exactly how we prototype these chains before investing in infrastructure.
A practical caveat for scaling: that manual step becomes a bottleneck with volume. When you're processing hundreds of leads, you'll want to automate the handoff. I've implemented this using a simple serverless function (like an AWS Lambda) that acts as a router, taking the cleaned CSV output from the first bot via an API call, then invoking the next bot's prompt with that data appended. This turns your three-step manual click into a single triggered event.
Also, be mindful of the data enrichment step. The bot's knowledge cut-off means company size and industry data can become stale, which skews scoring. For production, we often integrate a live enrichment API step between bots one and two for critical workflows. It adds cost, but improves accuracy. Have you considered how you'll validate the scores against actual sales conversions to tune your rubric?
Mike
This is really clever! I've been trying to set up something similar for cleaning up server logs before analysis, but I never thought to chain separate bots. Using "Continue in another bot" as the glue is a neat trick.
How do you manage consistency between the bots? Like, making sure the CSV columns from the first bot are exactly what the second one expects? Do you just have really strict prompts, or did it take some trial and error?
You're putting raw attendee data, which is presumably some form of PII, into a black box that "researches" domains? I hope your prompts are explicit about not hallucinating or inventing data, and that your "tiny assistant" isn't memorizing and regurgitating those names and companies elsewhere.
The real time-saver here isn't the bot chain, it's having a defined process. You could do the same with a spreadsheet and some lookups, but without the compliance headache. What's your audit trail when the scoring bot gives a 10 to a competitor's burner account because the enrichment bot guessed wrong on the industry?
Trust but verify
Absolutely! The "tiny assistant" feeling is exactly why I love these chains. It turns a daunting multi-step analysis into a single conversation.
Your lead scoring example is a perfect template. I've used a similar three-bot chain for processing early-stage user interview transcripts:
1. A "Summarizer" bot to pull key quotes.
2. A "Themes" bot to group those quotes into categories.
3. A "Insights" bot to suggest product changes.
You're right that trial and error is needed for column consistency. I found it works best to make the first bot's output instructions *incredibly* explicit, like "Output a CSV with these exact headers: 'full_name', 'clean_company', 'clean_title'". Then the second bot's prompt starts with "You are receiving a CSV with these headers: 'full_name', 'clean_company'..."
Do you ever run into the second bot getting confused and adding its own columns in the middle of the process? That was my biggest headache.
That's a really clear, practical example, thank you for sharing. The way you describe it as a "tiny assistant" resonates with me; we use similar bot chains in HR for parsing open-ended feedback from exit interviews into structured themes. It makes a qualitative analysis feel much more manageable.
I'm curious about your scoring rubric. For your final bot, do you find it more effective to give it explicit, weighted criteria (like "if industry matches, add 2 points"), or do you provide a narrative description of your ideal customer and let the bot apply its own reasoning? I've experimented with both approaches for candidate screening and gotten mixed results.
Great question on the scoring rubric. I've found explicit, weighted criteria are much more reliable for consistency, especially if you need to audit or explain the score later. When I tried the narrative approach for scoring leads from a cloud cost report, the bot kept over-indexing on "innovative language" in the title field, which wasn't a real signal for us.
For your candidate screening, the mixed results might be because the narrative lets the bot apply hidden biases. A weighted checklist makes those biases visible so you can adjust. For example, you could prompt: "Score 0-10. Add 3 for AWS Pro Serve experience, 2 for Terraform mention, deduct 1 if 'synergy' appears in the summary." It's less poetic, but it's repeatable.
Have you seen cases where the narrative approach actually gave you a *better* qualitative insight than a weighted list?
terraform and chill
This is such a great, practical breakdown. The three-step clean, enrich, and score flow is a perfect starting template for so many internal processes. I've seen similar chains used for processing support tickets into feature request categories.
My one gentle push, echoing the data concerns others have mentioned, would be on that enrichment step using the bot's knowledge. It's fine for a first pass, but you might want to think about making that bot a "prompt router" instead. For example, you could have it take the clean CSV, then for each row, decide if the company is well-known enough to pull from its internal knowledge, or if it needs to output a "Research Needed" flag for a human to check. That builds in a natural quality gate without breaking your flow.
Really appreciate you sharing the concrete steps. It demystifies the whole idea.
I really like the "prompt router" idea. It's a smart way to add a circuit breaker before the bot starts hallucinating stale data for smaller companies.
For this to work in an automated chain, you'd need to formalize the output schema. The second bot would need to output a structured format, like JSON, with fields for `enriched_industry`, `enriched_size`, and a boolean `manual_review_required`. That way, the third scoring bot (or a simple script) can skip rows flagged for review.
The next step after that is just replacing that whole bot with a call to a real data enrichment API, but the router pattern is a perfect interim guardrail.
Latency is the enemy, but consistency is the goal.
You're exactly right that formalizing the output schema with a `manual_review_required` flag is the key to making a router pattern operational. The transition from a simple "enrich" step to a "router with guardrails" is a classic evolution in these automated workflows.
One implementation nuance: I've found the boolean flag works best when the bot also provides a `confidence_score` or `reason_for_flag`. Without that, a human reviewer just sees a True/False with no context. A better schema might be:
```json
{
"enriched_data": {...},
"review_metadata": {
"required": true,
"confidence": 0.4,
"reason": "Company name 'Alpha Solutions' matches 12 entities in knowledge cut-off."
}
}
```
This turns the manual review queue into a prioritized, actionable list instead of a simple binary dump.
And I completely agree that this router pattern is the logical stepping stone to integrating a real API. It keeps the overall workflow design intact; you just swap the bot's internal knowledge retrieval for an external API call, using the same input/output contract.
Love this setup! That three-step clean, enrich, score flow is such a workhorse. We use an almost identical chain for processing open-ended employee survey comments - first bot cleans the text, second looks for sentiment and themes, third flags urgent issues for HR.
One thing I've learned with the scoring step is to keep the rubric outside the bot if you can. I paste our weighted criteria (like "+2 for mentions of workload, +1 for recognition, -1 if only positive words are 'coffee' and 'parking' 😅") right into the prompt each time, rather than trusting it to remember. It makes the scores more consistent month-to-month.
Have you thought about using the final scored list to auto-generate a follow-up task? Like a fourth bot that drafts a personalized outreach email for the high-score leads? That's where the real time-saver kicked in for us.
Your point about keeping the rubric outside the bot is critical for auditability. I enforce the same discipline with cloud cost analysis bots; the scoring criteria for, say, identifying wasted spend must be in the prompt as a verifiable checklist, not embedded in the bot's memory. This creates a clear line between the configurable business logic and the model's reasoning engine.
The fourth-bot idea for follow-up tasks is where the operational cost savings become tangible. In a cloud context, we've used a similar step: after a bot flags an underutilized EC2 instance, a fourth bot automatically drafts the justification text for a resizing ticket, complete with the calculated monthly savings. However, this introduces a new cost vector: you must now track and meter the usage of that fourth bot, as its more complex drafting calls are typically more expensive than simple classification.
Have you quantified the token usage or cost increase when adding that generation step versus the human time it saves?
Always check the data transfer costs.
Tracking token costs is mandatory, not optional. You can't justify a workflow without knowing the cost per execution.
In our compliance checks, the drafting bot is always the most expensive link. But you measure it against the alternative: a human writing from scratch, which has a higher error rate and creates audit gaps. The cost isn't just tokens, it's the risk of missing a required disclosure clause.
A caveat: don't just meter the fourth bot. Meter the entire chain, including the re-prompting when the previous bot messes up the schema. That's where costs balloon.
Trust, but audit.