Hey folks! I've been running Shield Advanced for about six months now to protect our startup's API infrastructure, and overall I'm a huge fan of the automated mitigation and the SOC support. It's been a game-changer for our peace of mind.
However, I've recently been dealing with a persistent issue that I'm not sure is a gap in my configuration or a limitation of the service itself. We're seeing what looks like a low-and-slow application layer DDoS vector that Shield Advanced doesn't seem to fully catch, and our WAF (with the AWS Managed Rules) is struggling to keep up efficiently.
Here's the gist: we're getting hit with thousands of requests per minute, but each request comes from a different IP (looks like a massive proxy or botnet), and they're all requesting valid, but computationally expensive, endpoints. Each request individually looks legit—proper headers, no obvious malicious payloads—so the WAF isn't blocking them. But collectively, they're spiking our backend CPU and causing latency for real users. It feels like a resource exhaustion attack.
What I've tried so far:
* Reviewed Shield Advanced dashboards – it shows elevated traffic but hasn't declared it a "DDoS event" or triggered its advanced mitigations.
* Tweaked WAF rate-based rules, but with the IP rotation, the per-IP limits aren't effective.
* Enabled the 'Amazon IP Reputation List' and 'Known Bad Inputs' managed rule groups, which helped a little but not enough.
Has anyone else run into this scenario? I'm wondering:
1. Is there a specific way to tune Shield Advanced to recognize this pattern as a DDoS attack?
2. Should I be looking more at custom WAF rules based on aggregate request patterns (like total requests to a specific path across all IPs)? If so, any tips on setting that up?
3. Or is this just a type of attack that requires a different tool or approach altogether?
Really eager to hear if the community has any battle-tested strategies for this. The AWS docs are great, but real-world experience is priceless!
The behavior you're describing aligns with a known nuance in Shield Advanced's design. It's primarily engineered to detect and mitigate volumetric and state-exhaustion attacks at the network and transport layers (L3/L4). For a *low-and-slow* application layer attack comprised entirely of seemingly valid requests, the system often won't trigger a formal DDoS event declaration because, from a pure traffic volume perspective, it may not cross the algorithmic thresholds.
The key here is that Shield Advanced and WAF are complementary but distinct. Shield isn't a substitute for granular application logic control. Your mitigation for this vector will need to be driven by WAF rules and potentially application-level rate limiting. Have you explored implementing a custom WAF rule based on a aggregate count of requests to those expensive endpoints, say over a five-minute period, regardless of the source IP? This moves the detection from the individual request to the collective pattern Shield is missing.
Every dollar counts.
Yeah, the "complementary but distinct" line is just marketing for "you bought two things, not one."
The real issue is that everyone expects one toggle to solve every layer. WAF custom rules are the answer, but good luck tuning them without breaking real users. It's a pattern recognition problem they're selling as a simple checkbox. You'll spend more time on the WAF rule logic and false positives than you ever did on Shield.
If it ain't broke, don't 'upgrade' it.
You've accurately described a challenging scenario. This is precisely the kind of attack that sits in the operational gap between network-layer DDoS protection and signature-based WAF rules, a nuance that isn't always clear in documentation.
The lack of a formal DDoS event declaration by Shield is telling; its algorithms are designed to identify traffic anomalies that indicate volumetric attacks, not the coordinated resource drain you're experiencing. Since each request is technically valid, the system's heuristics at that layer won't be triggered. Your issue is one of aggregate intent, not individual request malformation.
Moving forward requires shifting your mitigation strategy from detection of bad packets to governance of good ones. The solution will likely be a combination of stricter, application-aware rate limiting at the WAF or API Gateway level and possibly architectural changes to offload or cache those expensive endpoint computations. Have you considered implementing a cost-based rate limit, where endpoints are grouped by their relative backend resource consumption rather than just a simple request count?
Let's keep it constructive
Interesting that it hasn't triggered a DDoS event. That's the core of the product promise, isn't it? Their event declaration is a black box, and you've just found a whole class of attack that lives outside its heuristics. It's a pattern recognition failure they'll call a "feature gap."
This isn't a configuration problem. You're seeing the limitation of any system designed to spot traffic spikes when the attack is designed to look like a slow, distributed rise in "good" traffic. Shield's dashboards show the symptom but the algorithm won't call it a disease.
Data skeptic, not a data cynic.
Yep, that's the whole game. The event declaration isn't a technical output, it's a contractual one. No event, no SLA breach. The "black box" conveniently defines the problem away.
You pay for the magic algorithm. When the magic doesn't work, you're told you bought the wrong magic. It's a brilliant business model.
Your stack is too complicated.
I understand the frustration behind viewing the event declaration as a contractual rather than technical line. That's a very fair interpretation of how it can feel in a situation like this.
It's a design limitation, but not necessarily a cynical one. The heuristic thresholds have to be set somewhere to avoid constant false alarms, and attacks that deliberately operate below those lines exploit that gap. The business model criticism is interesting, but I think it often stems from a mismatch between the product's marketed scope and a customer's real-world need for layered, holistic security.
In this case, the "magic algorithm" failing to declare an event is a clear signal that the attack vector falls outside its intended protection layer. It's a prompt to engage the other tools in the stack, like custom WAF rules or application logic, to address what is fundamentally a different kind of problem.
—daniel
You're hitting the exact scenario that shows why relying on Shield's event declaration is flawed. You said it's not declaring a DDoS event, but you're still experiencing a DDoS effect - resource exhaustion.
The problem is your expectation. Shield's event algorithm is looking for traffic anomalies that threaten AWS infrastructure itself, not your application's specific resource constraints. Your backend CPU spiking from valid requests is your problem, not theirs, according to their design.
The real conversation you need to have isn't about configuration. It's whether you bought a tool to protect AWS's network or to protect your business logic from being abused. They're not the same thing.
Trust but verify.