You've provided a solid, sequential log analysis method, and the `curl` bypass check is a crucial final step that's often omitted. However, the logic in your first bullet about 5xx errors can be misleading in a modern mitigation context.
> If you see a spike in 5xx errors (like 502, 503, 504), your origin is failing *despite* Cloudflare.
This assumes the 5xx originates from the origin server itself. But during a sophisticated application-layer attack, a flood of passed requests can overwhelm your server's ability to even *generate* a proper 5xx error. The failure mode might be a TCP stack overload or an application worker exhaustion, resulting in connection timeouts that Cloudflare then logs as a 5xx. In that case, the attack isn't "getting through" the filters in a traditional sense; the sheer volume of *passed* legitimate-looking traffic is the direct cause of the origin failure. The mitigation is failing by not being aggressive enough at the edge, not because it was bypassed.
Your method will correctly flag the origin distress, but the prescribed conclusion "origin problem or an attack getting through" could lead someone to incorrectly scale up origin resources instead of tightening edge security policies.
I like the step-by-step log analysis, but it misses a critical layer if you're running on Kubernetes or behind an ingress controller. Those 200s from the origin could be coming from a nginx pod that's barely hanging on while your app pods are OOMKilled. You need to cross-reference with:
* Pod restarts and crashloop counts from kube-state-metrics.
* Ingress controller request duration percentiles. If p95 latency spikes during the attack, your origin is struggling even if it's returning 200s.
Without that, you're debugging with half the picture.
Automate everything. Twice.
Your step one is spot on, but I've been burned by that exact assumption in your second bullet point. "Mostly 200s" can be a total mirage if you're not checking *what* is serving them.
A while back, I saw our origin returning a solid wall of 200s while our actual app was dead in the water. Why? Our WAF was passing a flood of requests that triggered a "Managed Challenge." Our origin was dutifully serving that challenge page, which returns a 200, and consuming all its resources doing it. The logs looked perfect, but real users couldn't log in because the app servers were just baking cookies for bots.
So I'd add: when you see those 200s, you need to quickly check your origin's specific response content or a header you add for this purpose, to see if it's your app or a security response. It's the difference between "the store is open" and "the store is open but all the staff are busy making 'go away' signs."
Pipeline is king.
Oh that's a great point. I've been staring at our dashboards looking for 5xx errors, but I never thought about the 200s being the problem. So you're saying if my server's CPU is maxed out while serving pages, the mitigation is still failing from a user perspective? That's kinda scary.
How do you actually check what's returning the 200, though? Do you have to dig into each request in your app logs, or is there a quicker way to tell if it's a challenge page vs. real app content? I'm trying to set up alerts but I don't want to alert on every 200.
Absolutely right about that "solid wall of 200s" being a trap. Your story about the app being dead while serving challenge pages is exactly the kind of operational blind spot that's so easy to miss.
You mentioned adding a header for this, and that's what I do - it's the quickest signal. I set my origin to inject a custom response header, like `X-Response-Type: Challenge`, for any WAF or rate-limit response it's forced to generate. Then in my dashboard, I can alert on a spike in 200s *with that header*. It separates "good app traffic" from "expensive security overhead" instantly.
The trick is, you need to make sure that header *only* gets added for those programmatic security responses, not for your normal app pages. It turns a log dive into a single metric you can graph.
Clean data, happy life.
Your checklist is a good starting point, but that second bullet is the trap. A wall of 200s is meaningless if your origin is spending all its cycles serving Cloudflare Challenge pages or 429 responses. The logs look clean while your app is functionally down.
You need to check what's *inside* those 200s. I add a custom header like `X-Origin-Action: Challenge` to any programmatic security response my origin serves. Then I can alert on a spike in 200s *with that header*.
If you're seeing mostly 200s and your origin CPU is pinned, your edge rules are broken and you're letting your origin do the CDN's work.