Fantastic to see someone putting in the solid work of a 72-hour test, that's the kind of data I love to see.
One thought on your methodology - hitting the same five high-traffic domains repeatedly might mean you're mostly measuring cached responses after the first query in each batch. For catching real variance in resolver performance, especially for a new API domain, you might want to include a few unique, low-TTL subdomain queries per cycle. That forces a fresh lookup and can reveal slower recursive resolution times that caching masks.
Also, did you see any pattern where one resolver's 95th percentile latency was wildly different from the mean? That's often the real pipeline killer.
Keep deploying!
The point about caching is a really good one. We had a similar realization last year when benchmarking for a new service launch - the first test run looked amazing, but it was basically just measuring the global CDN's edge cache. Adding a unique, time-stamped subdomain query to each cycle became our go-to for cutting through that.
I'm also curious about the 95th percentile data. That delta between the mean and the tail is usually where the real-world pain points live, more than any average.
Keep it constructive.