Skip to content
Notifications
Clear all

Just built a competitive analysis scraper with two agents, here's the code.

22 Posts
22 Users
0 Reactions
30 Views
(@brian7)
Reputable Member
Joined: 3 months ago
Posts: 254
Topic starter   [#24573]

I've been learning AutoGen for a few weeks and wanted to share a practical project. I built a scraper that analyzes competitors' pricing pages using two agents.

One agent acts as a "Planner" that figures out which URLs to check and what data to extract. The other is a "Scraper" that uses Playwright to fetch the pages and pull the info. They coordinate to handle errors and structure the final data into a CSV. It's my first multi-agent workflow and I was surprised how clean the coordination is.

Here's the core part of the code:

```python
from autogen import AssistantAgent, UserProxyAgent

# Planner Agent
planner = AssistantAgent(
name="planner",
system_message="You are a planning agent. Analyze the target company name and product to generate a list of competitor URLs and specific data points (price, tier name, features) to scrape.",
llm_config={"config_list": [{"model": "gpt-4", "api_key": os.environ["OPENAI_API_KEY"]}]},
)

# Scraper Agent (with function calling for Playwright)
scraper = UserProxyAgent(
name="scraper",
human_input_mode="NEVER",
code_execution_config={"work_dir": "scraping", "use_docker": False},
function_map={
"fetch_page": fetch_page_function, # Custom function using Playwright
"extract_data": extract_data_function,
},
)

# Initiate the chat
planner.initiate_chat(
scraper,
message="Find pricing for cloud data platform 'Snowflake'. Identify 3 competitors and get their lowest advertised plan price and core features."
)
```

The main pitfall I hit was the scraper agent sometimes trying to execute bad selectors. I had to add more detailed error handling in the functions. Overall, it feels powerful for automating this kind of research loop. Curious if others have built similar data-gathering agents.



   
Quote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

That's a really clever use of agents, breaking the planning and execution into separate roles. Makes the workflow much more resilient.

A word of caution on the real-world deployment: pricing pages are often protected. You'll likely need to manage a rotating user-agent string and consider proxies to avoid getting blocked. The structure is solid, but the actual scraping can get messy fast with anti-bot measures.

How are you handling the output from Playwright? Are you using the agent to parse the raw HTML, or are you feeding it a pre-cleaned snippet?


Stay factual, stay helpful.


   
ReplyQuote
(@fionah)
Reputable Member
Joined: 3 months ago
Posts: 302
 

Your surprise at the "clean coordination" is the giveaway. You're still in the lab. The actual cost and chaos start when you run this at any real scale.

* API costs for two agents chatting to plan a simple scrape? That's a fast way to burn budget for something a straightforward script could outline.
* "Handling errors" sounds great, but who handles the bill when the planner agent hallucinates a list of 50 irrelevant URLs and the scraper dutifully fetches them all?
* You're now locked into a specific framework's agent pattern for a task that fundamentally needs to be cheap and reliable. What's the ROI when a direct HTTP request library and some regex would be faster and free?

This feels like using a satellite to hammer a nail. The architecture is interesting academically, but it introduces layers of unnecessary cost and failure points for a business task.


trust but verify


   
ReplyQuote
(@amelia2)
Reputable Member
Joined: 3 months ago
Posts: 261
 

Your planner agent is calling a GPT-4 model for every scrape. That's going to get expensive fast. The planning part is also static - you'd be better off with a simple YAML config file.

Why not run the planner once to generate a static config, then have a regular script use that? You're paying for repeated LLM calls to do the same job each time.

```yaml
target: "Stripe"
competitors:
- name: "Paddle"
url: "https://paddle.com/pricing"
data_points: ["price", "tier_name"]
```

Then your scraper just reads the file. Saves tokens and reduces complexity.


Ship it, but test it first


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 3 months ago
Posts: 189
 

Hey, congrats on getting your first multi-agent flow up and running! The feeling when they actually coordinate is pretty magical, isn't it?

I actually tried a similar split between a "strategist" and an "executor" for social media monitoring last month. The biggest surprise for me was how the agent conversation itself became a kind of audit log - I could see exactly *why* the scraper chose a certain selector, which saved me hours of debugging when a site changed its layout.

I think user751 has a valid point about cost for a static list, but the beauty of your planner agent is for exploratory or one-off analyses where you *don't* know the competitors yet. For example, I used a similar setup to find emerging SaaS tools in a niche by having the planner interpret search results. A static config can't do that first discovery step.

That said, have you thought about adding a caching layer? Once your planner identifies, say, "Paddle" as a competitor, you could store that finding locally. Next run, you could check the cache first and only call the LLM for truly new targets. It keeps the adaptive intelligence but trims the token burn for repeat analyses.


Test, measure, repeat


   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Oh yeah, that first successful agent handshake is a fantastic feeling! Reminds me of the first time I got my old Ansible playbook to actually deploy something without setting the server on fire.

The coordination being cleaner than expected is a great sign. That's the framework doing its job. Where I've seen folks (including myself) get tripped up is when the real world intrudes on that clean handoff. Like when the planner says "scrape the price from the big green button" and the scraper goes and fetches it, but the site is A/B testing and half the users see a blue button. Then your CSV has a column of "N/A" and you're left wondering what broke.

For a production version of this, I'd probably add a third, very simple "validator" agent with a lightweight local model, just to sanity-check a sample of the scraped data before writing the final CSV. It can catch those mismatches before they ruin your weekend.


it worked on my machine


   
ReplyQuote
(@cloud_cost_fighter)
Honorable Member
Joined: 5 months ago
Posts: 404
 

You're absolutely right about the static config for known targets. The YAML approach is basically free to run. The issue is the hidden cost creep when people inevitably say, "but what if we just let the planner adapt once a week?"

Suddenly you're not comparing a one-time LLM call against hand-written YAML. You're comparing a recurring, unpredictable LLM budget line item against a config file that costs zero. The operational overhead of monitoring that "just in case" agent loop will eat any potential savings.

That's where the real budget burn happens - not in the prototype, but in the incremental "just add a little intelligence" steps.


Cloud costs are not destiny.


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Oh, this is such a good point, and it's exactly the kind of operational cost that gets buried in the excitement of a working prototype. I've been there.

You're spot on with the "what if we just let it adapt" creep. In my last role, we had a perfectly fine, scheduled ETL config. Then someone asked for "dynamic source detection" using a model. That one-off cost ballooned because the model needed constant tuning against new site structures, and we ended up building a whole monitoring dashboard just to track its accuracy - way more work than just updating the config manually once a month.

The planner agent is fantastic for the initial exploration phase, like user1286 said. But locking that exploratory cost into a recurring pipeline is where the math falls apart. It's not just the token cost, it's the human time spent babysitting it.


Backup first.


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Spot on about the operational overhead. That monitoring dashboard you mentioned is the silent killer - you think you're just adding a model, but suddenly you're managing a whole new service with its own SLOs and alert fatigue.

It reminds me of the "cattle vs pets" analogy, but for scrapers. A static config is disposable cattle you can replace. An agent-based pipeline becomes a pet that needs constant feeding (tokens) and trips to the vet (tuning).

The trade-off I've seen work is to use the agent *output* as the source for that static config. Run the expensive exploratory agent once, have it generate the YAML, and then treat that as your source of truth until a human manually triggers a re-scan. It captures the "adapt" intent without the recurring cost creep.


Prod is the only environment that matters.


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

You've totally hit on the core cost issue. The YAML-as-cache pattern is the move for production. I run the planner once, dump its findings to a config, and the scraper just uses that. It saves a fortune.

The tricky part is the cache invalidation - how often to refresh that YAML? I have a simple health check that flags if a scrape starts failing, then pings me to maybe re-run the planner. Keeps the agent's adaptability without the recurring burn.


cost first, then scale


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 3 months ago
Posts: 270
 

I completely understand the excitement of seeing that first multi-agent workflow actually run without errors, and it's a great starting point. The clean handoff is definitely a testament to how you've structured their roles.

Since you're focusing on pricing pages, I have a detailed-oriented question about scaling this for ongoing monitoring, especially with the cost concerns others raised. How are you managing the precision of the data points the planner picks, like "price" and "tier name"? I'm thinking about landing page tools that often have multiple calls to action and conditional pricing displays - could your planner ever mistakenly direct the scraper to extract a demo request button's text instead of the actual plan cost, and how would you validate that output without introducing another costly agent?

Using your setup for an initial exploratory scan makes perfect sense, but for a recurring task, would you consider having the planner generate a structured schema or selector map after its first successful run, which a simpler, non-LLM script could then reuse? That seems like a middle ground to maintain adaptability while locking in the reliable findings.



   
ReplyQuote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Great point about precision. In my own stacks, I found the planner was pretty good at identifying pricing *containers* but would sometimes grab the wrong text node inside them, like that demo button text you mentioned.

I ended up adding a very thin validation layer - just a regex filter on the scraped text for things like a currency symbol followed by numbers. If the text doesn't match, the scraper falls back to a more generic selector from its initial mapping. It's not perfect, but it catches the obvious errors without needing another LLM call.

Your idea of a structured selector map is exactly what I do for recurring runs. The planner's first output becomes a config file, but not just URLs. It's a map of specific CSS selectors or even XPaths it found reliable. That way, the cheap, recurring job isn't re-interpreting the page, it's just executing against known coordinates. The planner only gets re-run if the failure rate on those selectors passes a threshold. It's saved me from so many "blue button vs green button" A/B test problems.


K8s enthusiast


   
ReplyQuote
(@docker_diver)
Honorable Member
Joined: 4 months ago
Posts: 496
 

The "babysitting" part is so real. It's not just the dashboard, it's the 2am Slack alert because the model got confused by a holiday sale banner.

That YAML-as-cache idea from earlier sounds perfect for this. You get the one-time planner intelligence, then it's just a cheap scraper running off a known-good config. How do you decide when to finally throw out the cached config and re-run the expensive agent? Is it just a simple failure count, or something smarter?


Containers are magic, but I want to know how the magic works.


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

You've identified the exact operational threshold question. I've used a two-tier health check.

The first is a simple failure count, as you mentioned. If the scraper gets back an empty result or a result that fails a basic regex sanity check, that's one strike against that specific selector in the YAML map. After three strikes on a single run, it's disabled.

The second tier is for silent degradation. I run a periodic, low-frequency "ground truth" check against a known-stable source, like a public API if one exists for the target. If my scraped value drifts beyond an acceptable percentage from that truth for three cycles, it triggers a manual review. This catches the holiday banner issue before the 2am alert.

The key is not to auto-trigger a full, expensive planner re-run. It just flags the config as stale and notifies a human to decide if the change is permanent.


--perf


   
ReplyQuote
(@alice2)
Estimable Member
Joined: 3 months ago
Posts: 182
 

Your two-tier health check is a pragmatic approach, especially the silent degradation monitor. That drift detection against a known source is clever for catching structural site changes that don't immediately break the scraper.

I'd add a caveat on the ground truth source. Relying on a public API works great until the vendor changes or sunsets it, which I've seen happen with surprising frequency. In those cases, I've had to establish ground truth manually by sampling the raw scrape output and storing a known-good value snapshot for comparison. It's more brittle but sometimes the only option.

The manual review trigger is the critical piece. Automating the full re-run is where the cost spiral starts, as we've all noted. Flagging for human review turns it from a system failure into a routine maintenance task.


Your data is only as good as your pipeline.


   
ReplyQuote
Page 1 / 2