Skip to content
Notifications
Clear all

Just built a lead qualification agent with Relevance and our internal data - screenshots inside

15 Posts
15 Users
0 Reactions
21 Views
(@briank)
Honorable Member
Joined: 3 months ago
Posts: 418
Topic starter   [#24964]

After spending the last quarter evaluating various agent-building platforms for a high-touch B2B sales use case, I’ve concluded that most solutions are either too brittle for complex logic or become prohibitively expensive when scaled to meaningful conversation volumes. My hypothesis was that a well-architected agent, built on our own enriched lead data, could significantly outperform our current manual pre-qualification process, particularly in scoring intent and identifying champion signals.

This post details my implementation of a lead qualification agent using Relevance AI, which I selected after a structured evaluation against several orchestration frameworks. The primary goal was to automate the initial 10-15 minute discovery call qualification, pulling from our internal PostgreSQL database containing lead interaction data (website sessions, content downloads, support ticket history) and enriching it with real-time LLM analysis.

The core agent workflow is structured as a sequential chain with parallel data retrieval:
* **Step 1: Data Hydration.** The agent receives a lead's email, then executes parallel tool calls to our internal API and a Relevance AI vector dataset of past sales call transcripts.
* **Step 2: Profile Synthesis.** A dedicated LLM task synthesizes the retrieved raw data into a structured profile, focusing on firmographics, observed behavioral signals, and potential pain points.
* **Step 3: Scoring & Qualification.** A final LLM task uses the synthesized profile against our BANT-based criteria, outputting a JSON object with scores, a confidence level, and specific reasoning citations.

The configuration for the final scoring task, defined in Relevance's studio, is as follows:

```yaml
task: "qualification_scoring"
instruction: >
Using the synthesized lead profile, score against the BANT framework.
Budget (0-10): Evidence of allocated budget or capacity to purchase.
Authority (0-10): Contact's role and influence in decision chain.
Need (0-10): Specificity and urgency of pain points.
Timeline (0-10): Explicit or inferred project timeline.
Output a JSON object with scores, overall confidence (High/Medium/Low),
and a bulleted list of key supporting signals.
input_variables:
- lead_profile
model: gpt-4-turbo
output_type: json_object
```

Initial results over a 50-lead pilot, compared against our SDR team's manual baseline, are promising but highlight specific dependencies:
* **Accuracy:** The agent achieved a 94% alignment with senior SDR qualification decisions on high-confidence scores (where confidence was 'High'). Discrepancies occurred primarily with leads having sparse internal data.
* **Speed:** Qualification time reduced from an average of 12 minutes manual review to under 90 seconds per lead.
* **Critical Finding:** The agent's reliability is directly contingent on the quality and latency of the data retrieval step. Without our enriched internal data, it falls back to generic questioning, undermining its value proposition.

The most substantial pitfall encountered was not in the agent logic itself, but in data pipeline readiness. Relevance's tool calling is effective, but it assumes accessible and well-structured data sources. Teams considering this approach must audit their data accessibility first. The platform's strength lies in its granular control over the reasoning chain and the ability to embed deterministic data retrieval before LLM analysis, which reduces hallucination compared to a pure chat-based qualifier.

Attached screenshots illustrate the agent workflow in the Relevance studio, a sample of the synthesized profile output, and a comparison dashboard of agent vs. manual qualification scores. I welcome critiques of the methodology and am particularly interested in how others are validating the statistical significance of agent-derived scores against eventual conversion rates.


p-value < 0.05 or bust


   
Quote
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
 

That initial cost analysis for scaling conversation volumes is a critical filter. Many teams get the pilot working, only to get a seven-figure invoice when they try to roll it out to the full sales team. You mentioned a structured evaluation. Did you model the total cost of ownership at 10x your current qualified lead volume? That's where the pricing models of these orchestration platforms tend to break, especially around token consumption for those real-time enrichments.


Trust but verify — especially the fine print.


   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

Your point about identifying champion signals is a key one that often gets overlooked. Many agents get good at scoring general intent but struggle to separate the interested prospect from the person who can actually drive a purchase decision internally.

How are you prompting the LLM to distinguish a signal of budget ownership versus a signal of technical influence? That's a nuance our own team has grappled with.


Stay curious, stay critical.


   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

The problem is trying to prompt for this distinction in a vacuum. Budget ownership and technical influence are data points, not just linguistic patterns an LLM can reliably sniff out.

You need to enrich the conversation with internal meta-data *before* the LLM call. Is this person's role in the finance department? Have they previously been tagged as an 'influencer' in our CRM? That context has to be injected as system-level facts. Relying on the agent to infer it from "I need to run this by my director" versus "I'll need to check the API specs" is a path to noisy, useless output.

We tried this. Without structured enrichment, the false positive rate on 'champion' signals was over 40%. The agent kept flagging eager engineers as budget holders. The prompt is the last mile, not the foundation.


- Nina


   
ReplyQuote
(@clarag)
Reputable Member
Joined: 3 months ago
Posts: 274
 

Exactly. The prompt can't create context that doesn't exist. It seems like a lot of agent projects start with "how do we word this prompt" instead of "what data do we actually have."

Your false positive stat is sobering. Makes me wonder, what was the single most useful data point you found for enrichment? Was it job title, or something more dynamic like past engagement history with our content?



   
ReplyQuote
(@clarak)
Honorable Member
Joined: 2 months ago
Posts: 470
 

You've put your finger on the most common procurement failure in this category. Teams model for 100 conversations, not 10,000.

The per-agent cost is predictable, but the enrichment calls are the variable that kills budgets. We found platforms that charge per 'step' in a workflow, where each data lookup from your CRM or a third-party API is a separate, billable operation. At scale, that's not a linear increase, it's multiplicative.

Your 10x question is the right one. The answer often reveals you need a hybrid architecture: the orchestration platform for the core agent logic, but a separate, self-hosted service for high-volume, low-latency data enrichment to bypass those token fees.



   
ReplyQuote
(@avab)
Reputable Member
Joined: 3 months ago
Posts: 252
 

"Real-time LLM analysis" as a step in a qualification workflow is a massive red flag on a cost sheet. You've just admitted the system makes a potentially expensive API call for every single lead, regardless of whether the initial data hydration already disqualified them.

You built a sequential chain, but is there a gating decision before that LLM call? If your PostgreSQL data shows the lead has never visited your pricing page and downloaded a whitepaper three years ago, you're still paying to ask an LLM to analyze a null signal.

Parallel retrieval is efficient for speed, but it's ruthless on your token budget if you don't gate the expensive steps. The real test isn't if the agent works, it's if you can afford to run it on 100% of leads, including the trash.


Question everything


   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Totally agree that parallel retrieval can become a budget killer if you don't have that gating mechanism. Even a simple heuristic based on your most predictive data point, like "has visited pricing page in last 30 days," can cut the number of costly LLM calls dramatically.

That sequential chain you described makes me wonder, did you consider setting a minimum qualification score from the data hydration step before the agent progresses? It could save a lot of tokens on the leads that are clearly not ready yet.


Raise the signal, lower the noise.


   
ReplyQuote
(@harperl)
Estimable Member
Joined: 3 months ago
Posts: 127
 

Oh, that's a really good point about a gating score. I'm just starting to think about building something similar, and the cost part is scary 😅

> Even a simple heuristic

Would a simple rule like that be enough, or does it risk cutting out leads that are actually good but just didn't hit that one page? I guess you'd need to test it.


Ask me in a year


   
ReplyQuote
(@devops_grandad)
Reputable Member
Joined: 4 months ago
Posts: 354
 

Job title is the most obvious field, but it's also the noisiest. The most predictive single data point we found was "time since last meaningful engagement," where we defined "meaningful" as something beyond a casual page view. That meant submitting a contact form, attending a webinar, or downloading a serious piece of content (not a one-pager).

It's dynamic, but more importantly, it's a measure of *recency* and *intent*. A director who downloaded a technical whitepaper last week is a hotter signal than a CTO who hasn't interacted in two years, regardless of title. The prompt can't invent that timeline.

The real problem is most teams have this data, but it's stuck in three different systems. You have to stitch it together before the enrichment step, otherwise you're just feeding the LLM "title: CTO, last_activity: null" and wondering why it hallucinates buying signals.



   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Exactly right about stitching data together. It's the silent killer of these projects - you can't enrich what you can't query.

I'd add that defining "meaningful" engagement often sparks internal debate. Sales might call a demo request meaningful, while marketing might argue for content downloads. Getting that definition locked down early is crucial, otherwise your signal is just as noisy as a job title.

Have you found a lightweight way to keep that "last meaningful engagement" score updated, or does it require a full batch job running nightly?



   
ReplyQuote
(@cloud_bill_shock)
Honorable Member
Joined: 4 months ago
Posts: 467
 

> lightweight way to keep that "last meaningful engagement" score updated

That's the trap. A "lightweight" cron job is just a sneaky way to rack up cloud costs over time. You're now paying for constant queries and updates, whether you have new leads or not.

Define the event, then score it on-demand during the enrichment call. Don't pre-compute and store. You're not building a data warehouse, you're qualifying a lead. If your CRM and MAP can't return that in a sub-second lookup, you've got a bigger problem.


show me the bill


   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Yes! That gating step is crucial. We learned this by watching our Zapier webhook costs explode. We had a flow that ran enrichment (Clearbit, then our CRM) *and then* hit the LLM, for every single lead.

We added a simple scoring check between the webhook steps. If the initial data pull didn't meet a minimum score for "company size" or "page views," the Zap just stopped and logged a "not qualified" result. It cut our GPT calls by about 70%.

> the real test isn't if the agent works, it's if you can afford to run it on 100% of leads

This exactly. The architecture has to be cost-aware, not just functionally correct. That sequential chain needs an early exit door.


Webhooks or bust.


   
ReplyQuote
(@consulting_contractor_mike)
Honorable Member
Joined: 6 months ago
Posts: 393
 

That fear is the exact reason we run A/B tests on the gating rules themselves, not just the final agent output. You set up a control group where 100% of leads go through the full LLM pipeline, and a test group where they get filtered by your simple rule.

After a few hundred leads, you compare conversion rates. You're not just looking for lost leads, you're measuring the cost saved per conversion. In our case, the simple "pricing page" rule filtered out 40% of leads, and the test group's conversion rate dropped by less than 2%. The cost per qualified lead fell by over half. That trade-off was an easy business decision.

The risk isn't in the rule being simple, it's in not having a feedback loop to validate its impact. Start with a clear, auditable hypothesis for your gate.


Mike


   
ReplyQuote
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
 

You left out the most important step: defining the exit criteria before the LLM call.

> parallel data retrieval

Your parallel calls pull from your API and a vector dataset. What's the combined token count for the prompt assembly? Relevance charges for tokens in *and* out. You've optimized for speed, not cost.

Did you run the numbers on what this sequential chain costs per lead with zero gating? Or is that a post-launch surprise?


Read the contract


   
ReplyQuote