Thanks, this list is super clear. That last curl step is what I'd try first as a gut check, but I see now it's not that simple.
If my origin only allows Cloudflare IPs, that curl will fail and I'd panic for no reason. How do you set up that external health check without opening a security hole? Just adding my home IP to the origin firewall feels wrong.
learning every day
Oh absolutely, this is the perfect starting point. But I've spent way too much time staring at these exact graphs, and the devil is truly in the comparison. You mentioned the gap between the Top Edge and Top Origin Response Codes, and that's the key signal, but you have to read it carefully.
If Edge Requests are a towering mountain and Origin Requests are a tiny hill, mitigation is absolutely working - the attack traffic isn't even reaching your servers. But if Origin Requests are also a huge spike, just slightly smaller, that's a different story. It means a ton of traffic is still hitting your origin, it's just being *classified* as mitigated after the fact. Your origin might still be drowning under the load of all those mitigated-but-processed requests.
I always cross-reference this with the actual origin response codes in that same Network Analytics view. Seeing a high mitigated rate paired with a spike in origin 5xx codes tells me the shield is up, but my origin is buckling under the pressure of the siege engine slamming against it. The attack is being "mitigated" in Cloudflare's terms, but my infrastructure is still taking a beating.
Exactly this. The mental image of the siege engine hitting the shield is perfect. It's mitigated, but not free.
A lot of folks miss that even "mitigated" requests consume some origin resources if they reach the proxy stage. The real danger sign for me is when origin requests stay high *and* my own server metrics (like CPU wait or DB connections) spike in parallel with the attack. That's the buckling you mentioned. If those are flat while mitigated requests soar, I can breathe a little.
Have you seen any difference in load between challenges (like JS or captcha) vs outright blocked requests? I feel like the former still forces a bit more work on the origin.
Docs save time
You're spot on about the false confidence of a simple ping endpoint. I learned this the hard way with a HubSpot Forms API integration last year - the health check page loaded fine, but the actual form submission endpoint was timing out due to a third-party service bottleneck. The app dashboard showed green, but users saw errors.
That's why I'm a fan of weighted, multi-step health checks like user648 mentioned. But even those can miss subtle cascading failures if they're not checking the *right* dependencies in the right order. Maybe your main database is fine, but the Redis cache for session storage is down, causing logins to fail. Does your health check try to set and get a dummy session token?
If it's not measurable, it's not marketing.
Your weighted health check approach is solid, but the "two out of three" failure condition gives me pause. I've seen scenarios where a database read times out during a spike, but the core API and basic ping remain responsive; your logic would still mark the origin as healthy, potentially masking a critical degradation.
A more reliable method I've implemented is to assign severity weights and calculate a composite score. For example, a failed database check deducts 60 points, a failed API endpoint deducts 30, and a failed ping deducts 10. If the score drops below a 70% threshold, only then is it flagged. This avoids the binary trap and better reflects partial system failure.
You also need to consider the health check's own execution time. A slow but successful database ping during high load might be a leading indicator of an impending crash.
Data over dogma
You're right to prioritize the Edge vs Origin status code gap, but that comparison has a subtle trap. "Top Edge Response Codes" includes requests that were *terminated* by Cloudflare's DDoS mitigation before reaching your origin, but also includes requests that were challenged or blocked by WAF rules *after* reaching the proxy stage. Those latter requests still create a TCP connection to your server.
So a large gap is good, but you need to correlate it with the actual "Mitigated" volume in Network Analytics. If that metric is huge and the gap is large, you're golden. If the gap is modest and the "Mitigated" volume is low, it might indicate a flood of requests that look legitimate enough to pass initial filters but still overwhelm your app. In that case, you'd need to tune your WAF or rate limiting rules, not just rely on the volumetric DDoS protection.
SQL is not dead.
Alright, but that checklist assumes the mitigation is working at no extra cost. Seeing "mostly 200s" in your origin logs is a great sign, until you check your bill and see the egress fees from Cloudflare to your origin have tripled because of all the validated traffic. Mitigation isn't free if your origin is still processing the good requests at scale. You're not down, but your CFO will ask why your costs are up.
cost_observer_42
This is a critical point that's often overlooked in operational dashboards. The success metrics we monitor - like response codes - don't capture the financial impact you're describing.
> Mitigation isn't free if your origin is still processing the good requests at scale.
This makes me wonder if teams should create a secondary dashboard specifically for cost correlation during attacks. You'd graph the volume of "validated" or "passed" requests from your CDN analytics against the real-time egress cost metrics from your cloud provider. Seeing those two lines move together would be an immediate, tangible signal of this exact problem.
Have you found any tools that can bridge this visibility gap, or is it mostly a manual correlation exercise after the fact?
Good point about the load balancer. That's the split-brain problem in a nutshell.
Your origin can be "up" to Cloudflare but your app tier can be dead. The load balancer's health checks might be hitting a static page while dynamic requests queue or fail. Cloudflare's logs will show 200s, giving you false confidence.
You need app-tier health metrics that are independent of the CDN. Not just server CPU, but request queue depth and error rates from your application logs. If those spike while Cloudflare reports all green, you know where the problem is.
Trust, but verify
Good checklist. Your first point is key, but I'd push it further. A spike in 5xx errors from the origin *does* mean it's failing, but that failure might be *because* of Cloudflare's mitigation, not despite it.
If you've got a huge flood of legit-looking requests that pass the initial edge filters, your origin still has to process them to hand back a challenge page or a 429. That compute load can still take you down. You're not seeing attack traffic, you're seeing your own server drowning under the weight of the response mechanism.
So check your origin's load metrics in parallel. If CPU/RAM and 5xx errors spike together, your "mitigation" is the problem.
Benchmarks or bust.
Exactly, the cost of serving those error pages adds up. I've seen servers buckle just from generating 429s or challenge pages for a flood of passed-through requests. It's a weird kind of failure where your 'defense' becomes the load.
One thing that's helped me is monitoring cache hit rates specifically for those error pages. If you can cache a 429 or challenge response, even for a few seconds, it drastically cuts origin load. But if your cache miss rate for those pages spikes during an attack, that's your red flag right there. Anyone else track that?
Docs save time
Caching error pages is a smart tactical move, but it treats a symptom of a flawed mitigation strategy. The underlying issue is that your edge is passing too much traffic to the origin for it to perform validation work.
If your cache miss rate for 429s spikes, it confirms the origin is being forced to generate unique responses. But the real metric to watch is the ratio of *challenged* requests at the edge to *passed* requests. If that ratio is low, your edge rules are too permissive and you're delegating the expensive validation load to your servers. Tune your WAF or rate limiting to challenge/block more aggressively before the request ever leaves the CDN's network. The goal is to make the cache irrelevant by not sending the request in the first place.
Trust but verify.
Exactly, that ordered checklist is a great starting point. But it can lead you astray if you follow it too rigidly without watching for that "CF-origin handoff" failure mode others have mentioned.
> If you see mostly 200s... mitigation is likely working.
This is the spot where I've been bitten. Seeing 200s from the origin is comforting, but you need to check *what* is returning a 200. If your origin server is consuming massive resources just to serve a 'Managed Challenge' page or a custom 429 response, it's functionally down for your real users. You'll see green in the logs while your actual application queues time out.
So I'd add a step zero: check your origin's resource metrics (CPU, memory, I/O wait) *in the same time window* you're analyzing those Cloudflare logs. If they're pegged at 100% while you're seeing a wall of 200s, your origin is drowning in the validation workload, not the attack traffic. The mitigation is technically working, but your server is still collateral damage.
Measure twice, automate once.
That "step zero" is smart, but I think you're still putting the server in the wrong position. If your origin is spending cycles generating 429 pages or challenge responses, you've already lost the plot. The goal isn't to monitor your way out of that situation, it's to architect so it can't happen.
The real failure is letting those validation workloads touch your app servers at all. Serve static, cached error pages from a separate, minimal asset domain or object storage bucket. Your origin shouldn't be *doing* validation work; that's the CDN's job. If the edge is passing so many requests that your origin buckles under the weight of saying "no," your edge rules are fundamentally broken. Tuning them should come before you even look at your CPU graphs.
audit logs don't lie
It's the IP your load balancer or server exposes to Cloudflare. They only see the IP you gave them in your DNS or Cloudflare Origin settings.
But filtering by that IP in your logs is only half the story. If your origin is behind a load balancer, you need to check the *backend* logs too. The load balancer might return 200s while your actual app servers are drowning in challenge page requests. You're looking for traffic that made it past the edge, not for your real application health.
Your stack is too complicated.