Everyone parrots the same line: "Use Tool X for hreflang audits, it's the industry standard." Then you actually run a large-scale migration for a global brand and watch those same tools light up like a Christmas tree with thousands of false positives, completely missing the structural errors that actually matter.
I've spent the last month neck-deep in this, auditing a site with 15 language variants and 50 country targets. The usual suspects—let's call them ScreamingFrog, Sitebulb, and a certain "enterprise" platform—all failed in spectacular but different ways.
ScreamingFrog is decent for catching missing return links, but its handling of x-default is naive. It flagged correct implementations as errors because it expects a self-referential link in every cluster, which the spec doesn't require. You end up wasting hours verifying non-issues.
The bigger issue is data staleness and crawl depth. Most tools use a live crawl for the audit, which is fine for a brochure site. For a dynamic site with geo-redirects based on IP headers? Useless. You need to simulate crawls with specific `Accept-Language` and `X-Default-Country` headers, which most GUI tools simply don't do. You're forced to script it.
```bash
# Example: Crawling the same URL for different hreflang targets
curl -H "Accept-Language: es-ES" https://example.com/product
curl -H "Accept-Language: es-MX" https://example.com/product
```
If you aren't doing this, your audit is a guess. The "enterprise" platform we tested cached a single version of the page from its last generic crawl and reported all hreflang tags as present, even though our CDN was serving entirely different HTML to users in Spain versus Mexico. The report was pristine, and completely wrong.
Then there's the issue of volume inflation. One platform reported "12,000 hreflang errors." About 11,500 of those were duplicate warnings for the same missing return link across paginated series, because it treats every page URL as a unique error, not every *hreflang cluster*. The signal is drowned in noise.
So, has anyone found a tool or method that actually works at scale, or are we all just writing custom scripts and pretending the shiny tools are doing the job?