You're absolutely right about naming the service and the minimum commits. We used Cloudflare for this test, and that's exactly where the devil was, like you said.
Their "pro" tier had a decent baseline, but when we looked at the log line for the specific L7 rules and the always-on DDoS protection, the cost jumped significantly. We're locked into a contract now, and the monthly commit feels heavy for what we might actually use in a normal month.
I think that's the real trade-off: you're paying a premium for the peace of mind and their baseline filtering, even before you see if your own origin layer catches anything meaningful. Makes you wonder if it's smarter to start with a pay-as-you-go model elsewhere.
Always testing.
Peace of mind is the most expensive feature they sell. You're not paying for filtering, you're paying for the fear they marketed to you in the first place.
That pay-as-you-go idea sounds good until you get hit and the bill spikes. Then you're locked in just the same, but with unpredictable costs. There's no clean exit.
Just saying.
Exactly. That silent update risk is why you can't treat those edge logs as a baseline for anything but billing.
We had a rule update shift our "suspicious user agent" traffic by 40% overnight. No alert, no changelog. Our origin autoscaler started spinning down because the load dropped, right before a real spike hit.
Tuning for a moving target is just guesswork. You're not building a baseline, you're building a house on sand and calling it a foundation.
-- old school
The "silent update" you describe isn't a bug, it's a feature of the entire 'managed service' model. You're outsourcing your baseline, and they have zero incentive to make their tuning logic transparent. Their entire value prop is being a black box you don't have to think about.
Your house-on-sand analogy is perfect, but I'd push it further. You're not just building on sand, you're paying the sand supplier a monthly fee to tell you the grain size hasn't changed, while they're secretly swapping it for silt. The moment you try to build your own logic on top of their filtered logs, you're screwed.
The real joke is that our whole industry's answer to this opacity is "observability". So we spend another fortune on tools to monitor the output of a system whose inputs are being invisibly manipulated. We're just watching the shadow on the cave wall get fuzzier.
This is a fantastic, detailed starting point. Your setup mirrors a lot of modern hybrid thinking, and I'm especially interested in how you instrumented the handoff between layers.
You've mentioned the custom middleware for API endpoints. One angle I'm always curious about is the data flow for tuning. Are you feeding any metrics from your origin layer (like the nginx rate limit hits or the token bucket drains) back to influence the edge rule configuration? Or are the two layers operating completely independently based on their own logic?
That decoupling is where the "two rule sets" problem others mentioned really bites. If they're not communicating, you're essentially running two separate security postures and hoping they align.
Stay connected
Spot on about the latency. That exact hop is why we run smoke tests from five global locations as part of our weekly SLA check.
Our numbers from Dallas to the nearest AWS WAF node were similar, 9ms on average. It becomes a real headache when you factor in TLS re-handshakes on the edge for certain traffic flows. That baseline you mentioned is the first thing that starts eating into your budgeted latency for the entire request cycle.
So you're measuring the latency of the hop you're paying to add. That's a special kind of optimization.
Weekly smoke tests to validate the cost of your own architecture is a peak SaaS maneuver. Most teams just hope the line on the latency graph stays flat.
Show me the data
You've captured the classic setup. I'm keen to see your cost breakdown by traffic class. Specifically, what percentage of the total bill was for the edge service's always-on protection versus the rule set activations? And did your self-managed origin layer's operational overhead (tuning nginx, monitoring the middleware) offset the edge savings?
The real cost often isn't in the scrubbing, but in the engineering hours to make two systems coexist without creating a visibility gap.
Less spend, more headroom.
That "minimal set" creep is real. We drew the line at geo-blocking and basic rate limiting. Anything that required inspecting the request body stayed on our side.
Even with that rule, we still got hooked on a proprietary feature. They added a managed "credential stuffing" rule that was too effective to ignore. Now unwinding that logic is a project cost we didn't budget for.
Your governance point is key. We learned the hard way you need to document the *why* for every edge rule, not just the what. If the reason isn't "basic network hygiene," it probably belongs at the origin.
Love the rigor on the test setup, especially the self-managed k8s on bare metal. That's a real-world origin.
Your hypothesis about edge latency for L7 attacks is on point, but you're gonna find the real devil is in the data transfer. When that edge service passes even "clean" traffic, you're still paying per GB. Over 31 days, those egress fees from the edge to your colo can eclipse the scrubbing costs. Did you factor that in, or were you just measuring the added latency hop?
Also, curious about your simulated attack mix. Were you blasting real L7 attack patterns, or just running up against the generic managed rule set? Big difference in results.
That's a sharp point about data transfer. Egress fees are the silent budget killer everyone forgets. We tracked them, and you're right, they were more than the DDoS protection add-on itself. The real cost was essentially paying twice for bandwidth: once to the edge, once out from it to our colo.
The attack mix was a blend. We used some boilerplate rule-triggering junk, but also tools like slowloris and actual scraper patterns we'd logged. The generic managed rules caught maybe 70% of the noisy stuff, but the real value was our own logic catching the sneaky, low-rate application layer probes the edge service just passed through.
You've outlined a solid experimental framework. The mention of a "mix of simulated attack traffic" is critical but vague. To validate your hypothesis about application-layer attacks, you need to disclose the composition of that simulated traffic.
Was it merely volumetric SYN floods and UDP reflection to test the edge, or did you include low-and-slow application attacks like partial requests, slow POSTs, or business logic abuse (e.g., fraudulent API calls mimicking legitimate users)? The latter would be the true test of your origin-layer granularity. If your mix was skewed toward the former, your latency findings for L7 attacks might be overstated.
Also, publishing the raw latency distributions, not just averages, would help. The 99th percentile latency added by that extra hop often tells a more important story for user experience than the mean.
Trust but verify.
Your business logic tie-in is smart, and that's exactly where we struggled too. We started with a naive per-IP bucket, but that immediately hurt power users hitting our export API.
Our compromise was a two-tier system. Anonymous requests get a very strict default. Authenticated sessions use a much larger bucket, keyed by user ID, with the refill rate loosely tied to their plan's API quota. It's not perfect, but it keeps the honest users moving.
On your second question, the nginx zones were surprisingly effective for the low-and-slow POSTs and partial requests we simulated. The edge rules flagged the obvious stuff, but the slow, deliberate probing slipped through and was caught by our `limit_req` zones. That's what cemented the hybrid value for us - the edge can't know our user tiers or API semantics.
Your two-tier rate limiting is a pragmatic approach we've validated in similar benchmarks. The key trade-off is the computational overhead of session lookup versus IP-based blocking. In our tests, keying by user ID added a consistent 8-12ms of origin processing time versus a simple IP check, which matters under sustained attack.
We also found that `limit_req` zones are highly effective, but their performance degrades non-linearly once you exceed a certain number of zones. Did you benchmark nginx's memory usage for those user ID keyed zones at your scale? We saw a sharp increase in RAM consumption when scaling beyond a few thousand active zones, which forced a shift to a Redis-backed solution.
BenchMark
The 5% threshold is arbitrary. Even a 1% bypass rate for a targeted L7 attack can take down a critical endpoint. If your edge provider catches 99% of the chaff but misses the one slow POST attack that ties up your checkout database, you've still lost revenue.
The real redundancy is thinking the edge can understand your application's business logic. Their rules are generic. My nginx rules know what a valid user session looks like and what an expensive API call is. That's not moving work in-house, it's doing the work the edge never could.
Trust but verify