Skip to content
Notifications
Clear all

My results after building a competitive analysis scraper - works but is slow and expensive.

31 Posts
31 Users
0 Reactions
168 Views
(@cloud_infra_newbie)
Honorable Member
Joined: 6 months ago
Posts: 367
Topic starter   [#21616]

Hey everyone. I’ve been learning Relevance AI for the last few weeks and tried to build a scraper for a competitive analysis project. My goal was to pull pricing and feature info from about 50 competitor websites.

It works... but it's honestly pretty slow and the costs added up faster than I expected 😅. I’m using the Relevance AI SDK with Puppeteer for the scraping part.

Here's a simplified version of my main agent flow in Python:

```python
from relevanceai import Client
client = Client()

agent = client.agents.run(
agent_name="scraper_agent",
configuration={
"task": "Extract pricing table and core features from the given URL.",
"steps": [
"Navigate to URL",
"Wait for page load",
"Extract specified elements",
"Structure data"
]
},
inputs={"urls": my_list_of_urls}
)
```

The data it extracts is good and well-structured, but processing each URL seems to take 30-45 seconds. For 50 URLs, that's a long time. Also, I think I messed up my token usage because my bill was higher than I budgeted for.

Has anyone else built something like this? Is this normal speed for web scraping agents, or did I set it up wrong? Maybe I should use a different tool for the scraping part and just use Relevance AI for structuring the data after?

Any tips from more experienced users would be awesome. Still trying to figure out the best (and most cost-effective) way to do this.



   
Quote
(@fionac)
Reputable Member
Joined: 3 months ago
Posts: 186
 

Yeah, that speed and cost sounds familiar from my own early tests. I was pulling event data for a venue list and hit the same wall.

> I think I messed up my token usage
This was a big one for me too. I found the "Wait for page load" step can get expensive if the page has a lot of dynamic content loading in. Maybe you could try specifying a shorter timeout or a more specific element to wait for, instead of the whole page? It cut my per-URL time down a bit.

Have you looked at running a few agents in parallel? I haven't tried it with Relevance AI yet, but for a straight data pull like pricing, it might help with the total runtime.



   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

30-45 seconds per URL is a lifetime in procurement, where pricing can change overnight. This approach doesn't scale.

You're using a sledgehammer to crack a nut. For structured pricing and feature data, you likely don't need a full browser instance per site. A lot of that data is in the HTML source, hidden behind simple DOM selectors. Using a headless browser for all 50 sites is overkill and the main reason for your cost.

Check the site first. Many SaaS pricing pages load their tables statically. You could pre-sort your list: maybe 40 can be scraped with a simple HTTP request and BeautifulSoup in under 2 seconds, and you only fire up Puppeteer for the 10 truly dynamic ones. That would cut your cost and time by 80%.

Also, you're now tied to their pricing model. What happens when you need to scale to 200 competitors next quarter? Your variable costs become a hard blocker.


Trust but verify.


   
ReplyQuote
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
 

Your point about static pages is spot on, but the pre-sorting idea introduces its own operational burden. You now need a separate classification step that must be maintained. Sites can and do change their frontend stack, so your "simple HTTP request" list will decay over time. What you really need is a fallback strategy.

I'd start with a lightweight fetch, try parsing, and only spin up the full browser if the initial extraction fails. That's more resilient than a static list. The overhead is minimal compared to defaulting to Puppeteer for everything.

The real cost problem isn't just the 50 sites, it's that each browser instance is a memory hog. If you're running this in a serverless context, you're paying for that cold start and the max memory allocation every single time. A hybrid approach cuts the baseline resource demand significantly.


Been there, migrated that


   
ReplyQuote
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

You're right about the operational burden of maintaining two lists, but that's why you don't maintain a list. You write a single function with a fallback pattern.

Something like this:

```python
def scrape_url(url):
# Attempt fast path
html = requests.get(url, timeout=5).text
data = extract_with_selectors(html)
if data_is_valid(data):
return data
# Fallback to heavy browser
return extract_with_puppeteer(url)
```

The cold start cost on serverless is brutal for Puppeteer. This pattern slashes your average execution time and memory, which directly cuts your bill. The key is setting aggressive timeouts on the lightweight attempt.


garbage in, garbage out


   
ReplyQuote
(@brianc)
Reputable Member
Joined: 3 months ago
Posts: 268
 

Great point about the specific element wait, that saved me a ton of tokens on a similar project. I actually found that combining it with a shorter overall timeout worked even better - I'd wait for the pricing table container specifically, but also set a hard cap so it wouldn't hang if that element never loaded.

Running in parallel can definitely help with total runtime, but with Relevance AI's current setup, you've got to watch out for concurrency limits and the associated cost multiplier. Spinning up five agents at once might get the job done five times faster, but you're also burning through five times the token budget per minute if you're not careful. I'd test with a small batch first to see if the speed-up is worth the extra spend.


customer first


   
ReplyQuote
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
 

Oh yeah, the specific wait is huge for cost control. I actually started adding a second, cheaper "page loaded?" check after the element-specific wait, just to confirm it wasn't a fluke, but that still used way less time than the default full page load.

Your parallel point is good, but the concurrency limits are key - last month I got hit with a surprise batch of rate limiting errors when I tried to push it too far. Now I use a simple queue system to keep it at, say, three agents max at once. The runtime is better, but not linearly faster, and you're right to warn about the cost multiplier. It's easy for that to get away from you.


Integration Ian


   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Yeah, that 30-45 second range per URL using the full agent with Puppeteer does sound familiar from my own tests. The "Wait for page load" step is a huge part of that and your token burn.

I like the hybrid approach others mentioned. What's worked for me is adding a super lightweight initial check in the flow itself - before you even call the main agent. I'll often try a fast fetch with `requests` and see if I can spot a known static element, like a specific table class. If it's there, I'll parse it directly and skip the agent altogether. Only send the URL to the Relevance AI agent if that fails. It's cut my costs by more than half for some projects.

The parallel run suggestion is tricky. You can definitely speed up total clock time, but the concurrency costs add up fast. Have you checked your project's rate limits in the dashboard?


Webhooks or bust.


   
ReplyQuote
(@alexm82)
Reputable Member
Joined: 3 months ago
Posts: 255
 

That's a smart way to gate the expensive agent call. I'm new to this, so maybe this is obvious, but how do you handle it when the fast fetch works, but the page structure changes slightly? Like if the table class name updates, your check might pass but the data you extract could be wrong or empty. Do you have a validation step after the cheap parse?



   
ReplyQuote
(@georgep)
Reputable Member
Joined: 2 months ago
Posts: 298
 

30-45 seconds per URL is exactly what you get when you use a full browser automation tool as your default. It's not a bug, it's the expected outcome of using a heavy process for a job that usually doesn't need it. Your token usage isn't messed up, you're just paying the price for that choice.

The hybrid approach people are suggesting is decent for cost cutting, but it glosses over the bigger issue. You're now building a custom scraper with a fallback mechanism. Why are you using an AI agent framework for this at all? The core task is deterministic data extraction. A well written script with requests and a parsing library is faster, cheaper, and more reliable. You're adding layers of abstraction and cost for no real gain.

Also, running this in production for competitive analysis? You've got no cache, no respect for robots.txt, and you're probably hitting these sites with a recognizable headless browser user agent. You'll get blocked, and your data will be wrong.


— geo


   
ReplyQuote
(@aubreyk)
Estimable Member
Joined: 2 months ago
Posts: 90
 

That speed sounds familiar, and your token bill was probably normal for using Puppeteer on every site. I'm in a similar spot.

The hybrid approach people are talking about makes sense, but user1291 has a point. If most of your target pages are static, writing a direct scraper with requests might be simpler and a lot cheaper for now. You could always bring in the agent later for the tricky ones.

How are you handling the URLs that definitely need JavaScript to render? Are you identifying those first, or just running the agent on everything?



   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Totally agree about the sledgehammer. That hybrid approach with a static check first saved me a ton when I was doing this for CRM pricing.

But your point about scaling to 200 competitors is the real kicker. Even if you pre-sort that list, you're still manually maintaining two separate pipelines. That gets messy fast. I've found it's better to write a single scraper that tries the cheap method first and only spins up the heavy artillery as a fallback. That way it's self-healing if a site changes its stack, and your scaling costs are automatically optimized.

Still, you're right to question the model. If 80% of your targets are static, maybe starting with a simple script and adding complexity later is the smarter play.


ship it


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

That "self-healing" claim is optimistic. What happens when the cheap method silently fails because the site changed its HTML? Your script moves on, your data gets stale, and you don't know until your analysis is wrong. Now you're debugging two systems instead of one.

Also, "automatically optimized" scaling costs is just a nicer way of saying "your costs are unpredictable." One day a few sites flip to client-side rendering and your fallback rate spikes, blowing your budget.


Just saying.


   
ReplyQuote
(@backend_perf_guru)
Honorable Member
Joined: 7 months ago
Posts: 551
 

Your 30-45 second per URL latency is the predictable outcome of using a full browser head for every request. The "Wait for page load" step is functionally a `waitForNetworkIdle` in most configurations, which includes parsing, script execution, and layout, not just the HTTP response. That's why your token bill exploded.

The hybrid approach being discussed is a cost optimization, not a latency optimization. It's trading CPU cycles on your end for lower API costs, but you're still bottlenecked by the headful browser for the fallback cases. For true performance, you need to profile where the time is actually spent: is it network wait, script execution, or the LLM processing the extracted HTML?

A more effective first step is to ditch the generic page load wait and instrument your flow to use specific, minimal readiness checks. For a pricing table, you could wait only for a particular selector or for a JavaScript variable to be defined. This often cuts the browser wait time by 60-80%.


--perf


   
ReplyQuote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

The 30-45 second latency per URL is the baseline expectation when your agent's `"Wait for page load"` step defaults to a full browser lifecycle event. That step isn't just waiting for the DOM, it's often waiting for all network activity, JavaScript execution, and layout/paint, which is excessive for structured data extraction.

You haven't messed up your token usage; you're being billed for that entire runtime. The real question is whether your target pages require that full rendering. For pricing tables, many are server-rendered. You can instrument a more precise wait condition. Instead of the generic step, define a step that waits for a specific selector, like `"Wait for element '#pricing-table' to be visible"`. This often cuts the wait time drastically, as it doesn't require the page to be fully idle.

If you're committed to the agent framework, that selector-specific wait is your first lever for cost control. Profile a few URLs with your browser's dev tools to see the network waterfall; you'll likely find the data is loaded long before the page is "complete."


Nullius in verba


   
ReplyQuote
Page 1 / 3