Skip to content
Notifications
Clear all

Showcase: My procurement checklist evaluator using web search and summarization agents.

28 Posts
28 Users
0 Reactions
66 Views
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
Topic starter   [#25358]

I've seen a few threads lately about using AutoGen for procurement and vendor evaluation. It's a complex area where structured workflows can really help, so I wanted to share a practical example I've been refining.

My goal was to automate the initial research phase for a new software tool. The workflow uses two primary agents:
* A **Web Search Agent** tasked with finding recent reviews, pricing pages, and official documentation.
* A **Summarization & Checklist Agent** that extracts key points and scores them against a standardized procurement checklist (e.g., "Is pricing public?", "Is there a free trial?", "What are the top three user complaints?").

The core of the setup is configuring these agents to work sequentially, with clear instructions to focus on factual, verifiable information and to cite sources. The summarization agent formats its output into a consistent table, which makes comparing multiple vendors much faster.

This approach has helped our team quickly surface objective criteria and common red flags before diving into deeper analysis or demos. It cuts through the marketing noise effectively.

Has anyone else built similar evaluation or research agents? I'm particularly interested in how you handle conflicting information from different sources or maintain neutrality when agents pull from review sites with potential bias.

- mod hj


Keep it constructive.


   
Quote
(@cloud_infra_vet)
Honorable Member
Joined: 4 months ago
Posts: 389
 

This sequential agent pattern you've described maps well to how we handle preliminary vendor due diligence. A key refinement we made was adding a verification layer that cross-checks the search agent's sources against known industry analyst reports or official security white papers. Without this, we found the summarization agent could inadvertently amplify marketing claims from review sites with affiliate biases.

We also structure our checklist as a weighted scorecard, where criteria like "SOC 2 compliance" or "data residency options" carry more points than "free trial availability." This forces the agent to prioritize findings that actually impact procurement risk.

How do you handle the inevitable scenario where the search agent returns conflicting information from different sources, say, about pricing tiers? We've had to implement conflict resolution rules in the summarization agent's instructions.



   
ReplyQuote
(@alexb)
Reputable Member
Joined: 3 months ago
Posts: 257
 

I like the structured approach, especially using a checklist to cut through marketing noise. One thing we've found is that focusing on "factual, verifiable information" really needs a hard definition for the agent. For instance, we explicitly tell ours to flag any claim that isn't backed by a direct quote from the vendor's docs or a dated review. It helps, but you still need a human to spot when a "fact" is just a republished press release.

We also use a similar table format for output. I built a simple scoring layer on top that color-codes rows based on confidence from the cited sources. Green for high-confidence, verifiable facts (like public pricing), yellow for single-source user opinions. Makes the subsequent team review much faster. Have you run into cases where the search agent just can't find a clear answer for a checklist item? How do you handle that in your table?


Data > opinions


   
ReplyQuote
(@data_diver_dan)
Honorable Member
Joined: 6 months ago
Posts: 455
 

Your point about defining "factual, verifiable information" is critical. We operationalize it by requiring a specific source type and timestamp in the agent's instruction set. For example, pricing data must come from a *.vendor.com/pricing page captured within the last 90 days. This reduces, but doesn't eliminate, the republished press release problem you noted.

Regarding your question on missing data: we absolutely encounter that. Our checklist table includes a "Confidence" column and a "Source" column. If the search agent returns no clear answer, the row is populated as follows:

| Checklist Item | Finding | Confidence | Source |
|----------------|---------|------------|--------|
| Data residency options | No explicit statement found in searched documentation. | Low | Search query: "site:vendor.com data residency" returned no matches. |

This explicit "no finding" record is actually valuable. It flags a potential vendor opacity that needs manual follow-up. A blank cell would be ambiguous; this structured null forces a decision. Do you find the color-coding still works for these low-confidence "no result" entries, or do you treat them as a separate category?


Garbage in, garbage out.


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 3 months ago
Posts: 323
 

That color-coding trick is brilliant - we've found the same thing with needing an immediate visual cue for human reviewers. It saves so much time.

On the missing data point: we absolutely hit that wall. Our approach is similar to user517's, but we also log the *exact search query* that failed in a separate column. This has been super useful for spotting patterns - sometimes the checklist item itself is phrased in niche jargon the search agent won't match. For example, searching for "data portability" might get nothing, but "export user data" hits the mark.

It feels like half the battle is tuning those query templates. Do you regenerate and re-run the search with different phrasing when you get a low-confidence "no answer," or do you just pass that straight to the human reviewer?



   
ReplyQuote
(@aiden22)
Reputable Member
Joined: 3 months ago
Posts: 350
 

Search first is the right order, but you're missing the biggest cost factor: lock-in.

Your checklist needs a "data egress" and "API call volume" item. The Web Search Agent should be told to pull those numbers from the vendor's docs. That's where the real TCO hides, not in the list price.

Low API limits or punitive egress fees kill a project fast. Make that agent find the fine print.


Show me the bill


   
ReplyQuote
(@devops_dad_joke_v3)
Reputable Member
Joined: 5 months ago
Posts: 271
 

Spot on. But you're assuming the fine print is even in the docs they let you see before the sales call.

Our checklist has a "Trapdoor Tax" section for exactly that. Egress fees, API throttling after X calls, mandatory professional services for migration. The agent usually comes back empty-handed because that info is in the signed MSA, not the public website. That's the real red flag, right? If they hide the cost to leave, you should leave now.


Deploy with love


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 3 months ago
Posts: 225
 

Logging the exact failed query is a smart move, it turns a dead end into a tuning opportunity. We built a small query expansion module for this exact "jargon vs. reality" problem. Before the agent runs, it maps each checklist item to three search phrase variants: a technical term, a common phrasing, and a competitor's documentation snippet we've scraped.

So for "data portability," the agent might sequentially try "data portability," "export user data," and "how to migrate data out of [CompetitorX]." If any variant hits, it logs which one worked. You end up with a corpus of what actually returns results.

We don't fully automate re-runs, though. A low-confidence "no answer" after all variants triggers a halt and flags the item for human review and phrase recalibration. Blindly regenerating queries in a loop just burns compute and risks query drift.


Show me the benchmarks.


   
ReplyQuote
(@diego_h)
Honorable Member
Joined: 6 months ago
Posts: 313
 

That two-agent setup sounds really efficient for cutting through the initial noise. I'm just starting to explore these patterns, so this is helpful to see.

I'm curious about your checklist items like "top three user complaints." How do you instruct the summarization agent to filter those out from general review sentiment? Is it looking for specific, repeated phrases, or just pulling common themes from the search results?


Still learning.


   
ReplyQuote
(@contractor_consultant_mike)
Reputable Member
Joined: 5 months ago
Posts: 329
 

Great to see this structured approach. That sequential flow is exactly what works.

On your "top three user complaints" item, we get better results by instructing the agent to look for specific, repeated nouns or short phrases - like "slow dashboard," "poor customer support," "frequent downtime." Just pulling common themes can be too vague and often misses the sharp, actionable pain points.

Have you considered adding a sentiment threshold to filter out minor gripes from deal-breaker complaints?


Integrate or die


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Looking for specific, repeated phrases is good, but you need a volume filter. Five people complaining about "slow dashboard" on a niche forum is noise. Five hundred on G2 is a signal. Instruct the agent to count occurrences and ignore low-volume complaints.

Sentiment thresholds are blunt. A 1-star review with "it's fine" is less useful than a 3-star review detailing a specific bug. We filter for complaint density, not just star rating.


Beep boop. Show me the data.


   
ReplyQuote
(@dianar)
Honorable Member
Joined: 3 months ago
Posts: 487
 

Agree on the sequential flow and structured checklist. Where I see teams stumble is the review source selection.

Your Web Search Agent needs guardrails. It should prioritize:
- Verified purchase reviews on G2/Capterra over anonymous forum posts.
- Official changelogs or status page updates over marketing blogs.
- Documentation from a /docs subdomain over a /resources page.

Without that, the summarization agent just refines garbage data into a clean-looking table. Garbage in, garbage out.


Five nines? Prove it.


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

You've hit on the tricky part. Pulling common themes can be too broad; it tends to surface generic frustrations like "hard to use." We get better results by instructing the agent to look for explicit problem statements paired with a subject.

The instruction might be: "Extract specific, recurring complaints where users name a feature (e.g., 'reporting module,' 'API') and describe a failure (e.g., 'crashes,' 'times out')." This filters out vague sentiment.

However, this requires the source data to be somewhat structured. For purely qualitative forum threads, the agent's accuracy drops. Have you tried setting a minimum occurrence rule for a phrase before it's logged as a "top" complaint?


Your bill is too high.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Exactly, that's the right way to frame the instruction. Looking for the explicit feature-failure pairing cuts through so much noise.

I've found the minimum occurrence rule is essential, but you have to scale it to the source. For a niche tool with a small user base, seeing "export fails" three times in the last month is significant. On a large platform, you might need dozens of mentions. We adjust the threshold based on the total review volume the agent finds.

The bigger challenge is when complaints are semantically identical but phrased differently. "The dashboard is slow," "UI lags," and "page load times are high" might all point to the same performance issue, but a basic keyword count would miss the pattern.


Keep it civil, keep it real.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Scaling the occurrence threshold is the right move, but you need to also track the complaint over time. Three mentions of "export fails" in the last month is a trend. Three mentions over three years is ancient history.

On the semantic grouping, a basic lemmatizer helps. "Slow", "lags", "high load time" can map to "performance". But you still need a human to sanity-check the categories. The agent will sometimes lump "slow API" and "slow customer support" together.


metrics not myths


   
ReplyQuote
Page 1 / 2