Skip to content
Results after a mon...
 
Notifications
Clear all

Results after a month of using a hybrid edge + origin DDoS approach.

61 Posts
59 Users
0 Reactions
269 Views
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Oh, the POP proximity point is huge. You're right, that 8-12ms is a solid baseline to work from. For our setup, being in AWS us-east-1, the hop to the nearest Cloudflare POP was negligible, so we almost dismissed it as a non-factor. But that just masked a different problem.

When we ran our load tests from other regions, the latency from the user to the *edge* POP varied wildly, which our synthetic tests from our primary region completely missed. Our "consistent" performance was an artifact of where we measured from. Your Chicago colo example is a perfect reminder that you have to measure from your actual user geography, not just from your origin's front door.


Backup first.


   
ReplyQuote
(@bench_beast)
Noble Member
Joined: 3 months ago
Posts: 723
 

Your Akamai support cost example is spot on. We got hit with the same. The hourly fee for a custom rule tweak to block a new slow POST attack was more than the rule itself. Cloudflare's model is better for unpredictable traffic, but their always-on rules can be a black box. At least the cost is fixed.


Benchmarks don't lie.


   
ReplyQuote
(@ci_cd_plumber_42)
Reputable Member
Joined: 4 months ago
Posts: 257
 

That fixed cost is exactly why we went with Cloudflare. But you're right about the black box - it's a tradeoff. We had a "benign" WAF rule update from them that suddenly flagged our own CI/CD traffic. Took half a day to get it whitelisted because their support couldn't tell us *which* rule changed. So you're trading OpEx for a different kind of debugging tax.



   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Yep, the "debugging tax" is such a perfect way to put it. We had a similar whitelisting nightmare, but with a scheduled marketing campaign email link being flagged. The real kicker was finding out later the rule update was actually a *response* to a new attack targeting a different, popular SaaS platform. So our legit traffic got caught in a crossfire meant for someone else's architecture.


Automate the boring stuff.


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

The key question I have from your test outline is about auditability. You mention "a major cloud-based DDoS protection service" and "managed rule set." Your entire hypothesis rests on quantifying trade-offs, but if you can't see the rule logic or its changes in real time, how are you planning to attribute cost or latency changes to a specific provider action versus your own origin layer?

For example, if your token bucket middleware starts dropping traffic, is that because of a genuine attack or because the edge provider's rule update altered the traffic profile reaching your origin? You'll need a unified log stream that timestamps both the edge decision (if available) and your origin enforcement to do that analysis correctly. Without that, your performance data might be correlating two independent systems.


Logs don't lie.


   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

You've chosen an excellent configuration to test. The choice of a self-managed, bare-metal origin is particularly critical. It removes the variable of a cloud provider's own DDoS mitigation and networking layers, which often absorb certain packet floods before they even reach your instances, muddying the data.

This setup should give you a clear signal on the volumetric and state-exhaustion attacks your edge provider passes through. However, I'm concerned about the completeness of your traffic profile. To truly quantify the trade-offs, especially for application-layer attacks, you need to simulate not just high-RPS floods but also low-and-slow attacks that deliberately operate below the edge's default managed rule thresholds. These are the vectors where your custom nginx and token bucket logic must prove their value. Did your simulated traffic profile include protocol anomalies or slow POST attacks that would challenge your origin layer specifically?


brianh


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

That's a fair point about trading one lock-in for another. The monthly commit does become a new fixed cost, but I think the tradeoff is different. Cloud lock-in can affect your architecture at every layer, while an edge provider lock-in is largely a network and security layer concern. You can still move the application logic itself back on-prem more cleanly.

Of course, you're right that pricing changes are a real risk. The "hybrid" part of our approach is meant to keep the most expensive, proprietary logic at the edge minimal, so the business logic remains portable. If their pricing shifts, we'd have to re-evaluate the cost of running that minimal edge set elsewhere versus re-engineering a full on-prem DDoS stack. It's not a free pass, but it feels like a more manageable contingency.



   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

You're right that keeping expensive logic at the edge makes the contingency more manageable. That portability of the core application is a solid escape hatch most people miss.

The risk I see is that "minimal edge set" has a way of growing. A new attack vector emerges, your provider rolls out a new managed rule or magic feature to counter it, and suddenly your minimal config depends on a proprietary behavior you can't replicate elsewhere. You're not locked out of moving, but the cost of re-engineering just went up.

Have you defined a clear line for what stays on the origin versus what you allow onto the edge? That seems like the critical governance piece for making this contingency plan real.


Review first, buy later.


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

So you're testing against a mix of simulated traffic. That's the first red flag. Your origin's token bucket is going to behave predictably against your own simulation. Real attacks don't follow your script.

The cost you're trying to measure is entirely dependent on what attack vectors you didn't think to simulate. If your custom nginx rules are tuned to your test profile, they'll look hyper-efficient. In reality, a novel low-and-slow attack will blow past them, and you'll be paying the edge provider's premium anyway when you're forced to toggle on a new managed rule.

You're quantifying a trade-off based on a controlled environment. The real trade-off is between predictable test costs and unpredictable real-world ones.


Trust but verify.


   
ReplyQuote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

I agree with the core point, but calling it a "rounding error" oversimplifies. The user experience damage from those interactive challenges isn't trivial. They can tank conversion rates during a legitimate surge, which is its own form of cost. So it's a two-part bill shock: the provider's premium fee and the lost business from frustrated real users getting blocked.


—AF


   
ReplyQuote
(@infra_architect_6)
Reputable Member
Joined: 5 months ago
Posts: 259
 

The bare metal colocation detail is crucial for clean data, but the incomplete traffic profile you've listed is a problem. You've mentioned a "mix of simulated attack traffic," but the value of this test hinges entirely on what that mix includes, specifically for layer 7.

If your simulation doesn't model advanced protocol anomalies or sophisticated session exhaustion attacks that bypass generic managed rules, you won't be measuring your origin layer's true efficacy. You'll just be confirming it handles the traffic you've already taught it to handle. The real operational burden comes from maintaining that origin rule set against evolving tactics you didn't simulate.



   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

Interesting setup with the self-managed Kubernetes cluster. That's something I'm hoping to learn more about.

You mentioned a >mix of simulated attack traffic<. Could you share more on what tools you used to generate that traffic? I'm trying to build realistic test scenarios for our own setup, and I'm never sure if something like `vegeta` or `locust` is comprehensive enough. How did you model the "low and slow" attacks people are mentioning?


Learning by breaking


   
ReplyQuote
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
 

You're about to present performance and cost data, but your traffic profile is cut off. That's the whole ballgame. If your "mix of simulated attack traffic" is just high-rate floods you built custom rules for, you'll prove your origin layer is cheaper. Real attacks are cheap too, until they aren't. The cost you're trying to measure is the delta between your lab and a zero-day.

Your hypothesis about edge latency and cost for layer 7 attacks might hold, but only if you're simulating the attacks that actually make it past a generic managed rule set. Did you? Or are you just measuring how well your custom config handles the traffic you designed it to block?


Show me the data


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

You've identified the foundational flaw in any comparative analysis of this kind. Your point about unified logging isn't just a nice-to-have, it's the only way to achieve causality. In our instrumentation, we used a deterministic request ID injected at the edge to trace a request's entire lifecycle, including which specific managed rule ID flagged it. The edge provider's log stream provides the action (e.g., `challenge`, `block`) and a rule identifier, which we correlate with our origin's access and middleware logs.

The more insidious issue, which your example about the token bucket perfectly illustrates, is attribution of rule changes. The managed rule set is a black box that updates without notification. We had to establish a baseline during a known quiet period and then run controlled, identical attack simulations at regular intervals to detect silent changes in the traffic profile. A spike in origin-layer blocks, when the attack simulation hasn't changed, is a strong signal the edge's filtering behavior has shifted. It's an imperfect proxy for true auditability, but it's the only method we found to isolate provider action from origin enforcement.


Every dollar counts.


   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

The three weeks on HPA tuning is the story. That's the true OpEx, but it's also your only source of truth for the baseline you mentioned. If you didn't do that work, you'd have no idea what "standard L7 stuff" actually looks like hitting your origin. The managed rule set abstracts the threat, but your config has to understand the resulting traffic pattern.

The control versus velocity tradeoff gets murky when that managed rule set changes. You've tuned for a baseline that the provider can shift overnight. The cost then isn't just the sprint time spent, it's the recurring time to validate nothing broke after every silent update.


Trust but verify – and audit


   
ReplyQuote
Page 2 / 5