Skip to content
Notifications
Clear all

Help: My website has 500k pages and every SEO crawler times out.

20 Posts
19 Users
0 Reactions
98 Views
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

Two passes is trading time for simplicity, which is valid. But this is where everyone's math goes off the rails.

You're spinning up and managing two full crawl systems. The compute and storage cost for the intermediate "clean URL list" at 500k entries isn't trivial if you're doing it right (storing metadata, status, maybe a hash). And you're still paying for the second crawl's execution time, which, even with a seed list, is a hefty resource block.

The real oversight is assuming the "polite delay" is free. Adding a 1-second delay between requests to be nice to your own server means that second crawl now takes 500,000 seconds minimum. That's nearly 6 days of continuous runtime. Your cloud instance cost for a 6-day continuous run, even a small one, will easily eclipse a month of many commercial tools. People forget to price their own time and the infrastructure carrying cost.


pay for what you use, not what you reserve


   
ReplyQuote
(@datadog_dave)
Honorable Member
Joined: 4 months ago
Posts: 494
 

Yep, the hidden limits are the worst part, aren't they? Been burned by that "silent sampling" myself. It renders the data useless for actual decision-making.

You're right about the self-hosted headaches, but there's a middle ground I've used successfully: treating the crawl as a pipeline. Use something lightweight like Scrapy for the discovery and fetch of all 500k URLs, dumping raw HTML to S3 or cloud storage. Then, you use a separate, scalable process (like Lambda or a batch job) for the actual analysis and auditing against that known set.

It splits the problem. The crawl becomes a known-quantity data collection job, and the heavy lifting (rendering, SEO checks) can be scaled independently without the vendors' black box. You still have to host it, but it's predictable and you own the data.


Dashboards or it didn't happen.


   
ReplyQuote
(@eval_rookie_42)
Honorable Member
Joined: 6 months ago
Posts: 445
 

That's really helpful to see the numbers laid out. I was wondering where that line was drawn.

If "large site" starts at 50k pages, then at 500k we're in a completely different league. Is that 50k limit consistent across most tools, or do some stretch further before the silent sampling kicks in?



   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

You're right about the data wrangling being a hidden time sink. That two-pass method with a filtered seed list is a solid way to get a clean dataset into a familiar tool.

But my caution with this approach is around that polite crawl delay. Setting it high enough to be "safe" can make the audit crawl take literal weeks. I've seen teams set a 2-second delay out of caution, then the project stalls because they're waiting over 11 days just for the fetch phase. It becomes a planning bottleneck.

The efficiency gain is real, but the clock is still ticking.


Trust the trial period.


   
ReplyQuote
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 404
 

You've hit the exact pain point. The silent sampling is what makes the data unusable, because you can't trust it for spotting site-wide patterns or rare issues.

At that scale, every tool I've tested has a hard ceiling before their algorithm decides what's "important" for you. Even some open-source tools you'd run yourself have default limits you have to reconfigure.

The only consistent workaround I've found is to not give them a full site to discover. Use their own API to feed a pre-compiled, batched URL list directly into the audit engine, bypassing their crawler entirely. It's more setup, but it forces them to process every page you specify.



   
ReplyQuote
Page 2 / 2