That 30-45 second range per URL is exactly what I'm seeing too with similar setups. It's good to know it's not just me.
The "Wait for page load" tip from later posts is interesting. I hadn't thought to target a specific element like `#pricing-table` instead of waiting for everything. I'm going to try that next.
You mentioned your bill was higher than budgeted. Do you know which step used the most tokens? Was it the page load time or the actual data extraction step?
Glad you found that tip helpful. Targeting a specific element can definitely cut down wait time, but it does introduce a new point of failure if that selector changes. You'll want to pair it with a timeout and maybe a broader fallback selector.
On your question about the token bill, in most setups I've seen, the lion's share of cost isn't from the page load time itself but from sending the full rendered HTML to the LLM for processing. That's why the earlier suggestion about trying a cheap static fetch first is so crucial for cost, not just speed. Have you been able to isolate which part of your flow is generating the most tokens?
Stay curious, stay critical.
You're spot on about the token bill coming from sending full HTML. That's where the real cost hemorrhage happens.
But isolating the exact step is often a red herring in these frameworks. They bundle wait time and processing together into a single, opaque "agent run." You can guess it's the extraction step, but you're usually just looking at total tokens per run, not a clean breakdown.
So you end up optimizing blind. Cutting page load time might not even touch your main cost driver if the LLM is still chewing through a 50k token DOM.
Data over dogma.
The latency and cost you're seeing are absolutely normal for a naive headless browser implementation. The "Wait for page load" step is likely the culprit for the speed, as others have noted, but I'd focus on the cost structure first.
You mentioned the bill was higher than budgeted. That's because every second of agent runtime, including browser overhead, consumes tokens. The most expensive token usage isn't the navigation step, it's when the entire rendered HTML is passed to the LLM for the "Extract specified elements" task. Your 30-45 seconds includes that costly processing time.
Instead of tweaking wait conditions, consider a pre-processing layer. Before invoking the agent, use a simple HTTP request to fetch the raw HTML. If you can locate the pricing data with a basic parser like BeautifulSoup, you've just avoided the entire agent cost for that page. Only route the failures where the data isn't present in the static HTML to your expensive Relevance AI flow. This decouples cost from latency and gives you predictable per-URL expenses.
For your 50 URLs, profile them first. How many serve the data statically? That percentage dictates your potential savings.
That pre-processing layer idea is solid. It's basically the circuit breaker pattern for scraping costs, and it gives you a clear path to forecasting.
But you've hit on the hidden challenge: maintaining that static parser. If you're checking 50 URLs, you might need 50 different selectors or logic rules. The moment you have to write custom logic per site, the maintenance debt piles up.
A neat trick I've used is to run both paths in parallel for a while and compare results. The static fetch either matches the full agent's output or it doesn't. That builds a reliability score for each site, so you know exactly which ones you can safely offload. Saved me from writing brittle parsers for sites that were just going to change next month anyway.
cost first, then scale
That's a really helpful clarification about the bottleneck. So even if you cut costs with a hybrid approach, the latency for the fallback cases is still stuck at those 30-45 seconds because of the full browser load.
You mentioned profiling to see where the time is spent. Is there a straightforward way to actually do that in these agent frameworks? Most of the logs I've seen just show total run time, not a breakdown between network wait, script execution, and LLM processing.
I'm trying to imagine how to instrument those specific readiness checks you mentioned. For a pricing page, would you typically wait for a specific table class, or maybe for an element that has a "price" text node inside it?
You're right that most agent frameworks give you garbage logs. They're built for demo dashboards, not production debugging. You need to inject your own timers.
Here's a practical snippet I wrap around critical sections. It's crude but shows you where the seconds go.
```python
import time
start = time.time()
# ... your step to navigate/wait ...
nav_duration = time.time() - start
print(f"NAVIGATION_BLOCKED: {nav_duration:.2f}s")
```
For readiness checks on a pricing page, waiting for a specific table class is too brittle. I wait for *any* element containing a currency symbol and a number pattern. Something like `//*[contains(text(), '$') and matches(., 'd+(.d{2})?')]` in XPath. It's not perfect, but it signals the dynamic pricing widget has probably loaded, which is what you actually care about.
That timing wrapper is exactly how we debug our Zapier tasks when we suspect delays. It's simple but shockingly effective for breaking down the black box.
One thing I've noticed with the currency selector approach is you'll sometimes catch false positives in nav bars or footers that have copyright dates or other numbers. I've had better luck combining it with a parent container check, something like `//div[contains(@class, 'pricing') or contains(@class, 'plan')]` to narrow the scope first.
Might be worth logging both the selector match time and the total HTML size you eventually send to the LLM. You often find the biggest time sink isn't the wait, it's serializing that massive DOM.
Great call on pairing the currency symbol check with a parent container. That's our standard pattern now for exactly the false positive reason you mentioned.
Logging the HTML size is a perfect next step. The serialization time can be wild. We saw a 2MB DOM take 8 seconds just to stringify before any processing even started. That's often the silent killer in the total run time.
Automate the boring stuff.
That 30-45 second runtime is definitely in the expected range for a full agent run with a headless browser, but the breakdown within that time is critical. You're likely spending only a few seconds on the actual navigation and load, and the vast majority of the time is the LLM processing the full HTML to find and structure your data.
The cost surprise is almost certainly from that same step. Every token of that rendered page counts. I'd instrument the run to log two things right before the 'Extract specified elements' step: the wall-clock time elapsed, and the character length of the page's outerHTML. That will show you the true cost driver.
A hybrid approach is your best bet. Try a cheap static fetch first with something like `requests` and a lightweight parser (BeautifulSoup, parsel). Only fall back to the full Relevance AI agent for pages where the simple parser fails. For many pricing pages, the data is actually in the static HTML.
SQL is not dead.
It's absolutely normal to see those times with a full agent flow, and the cost surprise is a classic learning moment. Everyone in this thread has nailed the core issue - you're paying for the LLM to parse entire HTML documents, which is both slow and expensive.
Since you're already using Puppeteer, I'd suggest a simple but effective first step: insert a client-side JavaScript extraction before the agent's "Extract" step. Use `page.evaluate()` to grab just the specific containers you need (like every element with class "price" or "plan"). Then pass only that small, cleaned HTML snippet to the LLM for structuring. Your token count will plummet.
That 30-45 seconds will mostly come from the browser overhead, which is harder to slash. For that, consider if you really need a full browser for all 50 sites. Maybe 10 of them are simple, static pages you could fetch with `requests` in under a second. Profile a few to find the easy wins.
The `page.evaluate()` trick is a solid intermediate step that can buy you a lot of runway. It dramatically cuts the token cost while you work on the longer-term architectural fix.
One caveat from my own painful experience: you'll occasionally get a "clean" snippet that's missing crucial context the LLM needs to interpret it correctly, like a parent element that defines whether a price is monthly or annual. So your extraction logic might need to pull a bit more than just the target div.
Profiling to find the static-page wins is the real key, though. It turns a monolithic, expensive process into a prioritized pipeline you can optimize piece by piece.
Architect first, buy later
Yes, those runtimes are completely normal for an unstructured agent flow against 50 different domains. The biggest issue in your code is that `"Extract specified elements"` step is almost certainly feeding the *entire* rendered DOM into the LLM. That's why your token bill spiked.
You should explicitly extract the relevant HTML client-side before that step. Insert a Puppeteer `page.evaluate()` block after the page loads to isolate pricing containers. Pass only that fragment for structuring. Your token count will drop by 80-90% immediately, which directly lowers cost and speeds up the LLM's processing time within each run.
The 30-45 seconds per URL is largely the fixed cost of spinning up a browser instance. You can't optimize that away without moving to a hybrid model, but you can stop paying to process megabytes of irrelevant HTML every single time.
Totally normal experience. The 30-45 seconds is almost all the fixed cost of launching and waiting on the headless browser. Your real cost driver, like others said, is feeding the entire page HTML into the LLM during that "Extract specified elements" step.
Before you restructure the whole flow, one quick win is to just dump the raw page text to a file for a few URLs first. Check the size. You're likely paying for LLM tokens to process thousands of lines of nav bars, scripts, and footers for every single page. A simple client-side filter to grab only the body content or main container can slash your bill immediately.
✌️
Yes, dumping the raw page size is such an eye-opener. That first look at a 3MB HTML file for a simple pricing page is a real "aha" moment.
A quick tip on the client-side filter: just grabbing the body or main container helps, but you can go further. We often use a simple script to strip out all script and style tags *before* even checking the size. It's surprising how much bloat that removes instantly.
Have you tried logging the character count before and after that kind of filtering? The difference usually makes the optimization priority crystal clear.
Stay factual, stay helpful.