Skip to content
Step-by-step: Setti...
 
Notifications
Clear all

Step-by-step: Setting up auto-scaling for origin under layer 7 attacks.

19 Posts
19 Users
0 Reactions
103 Views
(@cloud_cost_hawk)
Reputable Member
Joined: 3 months ago
Posts: 250
Topic starter   [#22268]

You're asking for trouble if your origin can't scale under a layer 7 flood. I see too many setups with a beefy, static origin server behind a WAF/CDN. When the attack bypasses the edge, your app dies and your AWS bill still spikes from the traffic transfer. The goal is to survive and keep costs from exploding.

Here's a practical, cost-aware setup using AWS primitives. This assumes you're already behind CloudFront or a similar proxy.

**Core Components:**
* **Application Load Balancer (ALB):** Non-negotiable. It's your scaling anchor.
* **Auto Scaling Group (ASG):** With scaling policies based on **RequestCountPerTarget**, not CPU.
* **EC2 Launch Template:** Using lean, pre-baked AMIs. No fat images.

**Critical Scaling Policy (CloudWatch Alarm):**
A layer 7 attack means HTTP requests. CPU might lag. Scale on the ALB metric.

```json
{
"AlarmName": "High-Request-Rate-Per-Instance",
"MetricName": "RequestCountPerTarget",
"Namespace": "AWS/ApplicationELB",
"Statistic": "Sum",
"Period": 60,
"EvaluationPeriods": 2,
"Threshold": 1000, // Tune this for your app's capacity
"ComparisonOperator": "GreaterThanThreshold",
"AlarmActions": ["arn:aws:autoscaling:...:policy/scale-out"]
}
```

**Cost-Saving Must-Dos:**
* **Use Spot Instances for the ASG mixed instances policy.** For a stateless app under attack, Spot handles the surge. Blend with On-Demand if you're risk-averse.
* **Set aggressive scale-in policies.** Once the attack stops, you don't need 100 instances idling. But use a longer cooldown to avoid thrashing.
* **ALB is expensive per LCU.** Monitor your `ActiveConnectionCount` and `NewConnectionCount` during attacks. If connections are high, consider scaling up instance size (fewer, larger instances) to reduce ALB cost, not just instance count.

**The Caveat:**
This keeps you alive, but it's still a financial hit. You're paying for the scaled-out EC2, ALB LCUs, and data transfer. The real solution is to stop the attack at the edge. This is your last line of defense. Pair this with WAF rate-based rules and geo-blocking to reduce the scaling load.


cost optimization, not cost cutting


   
Quote
(@austinm)
Estimable Member
Joined: 2 months ago
Posts: 123
 

Tuning that threshold is the real trick. You're right that CPU lags, but RequestCountPerTarget is useless if you don't know your actual per-instance RPS limit before things go sideways. That's a load test number, not a guess.

Have you priced this out during an actual sustained attack? The ALB cost itself becomes a massive multiplier. You might survive the app dying, but the bill shock from ALB LCU and data transfer while scaled out could be a different kind of outage.


trust but verify


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

The JSON snippet is cut off, but your core point stands: you must use the ALB metric. I'd push further and argue that even `RequestCountPerTarget` is too slow during a rapid-onset attack. You need a composite alarm.

Scale on that metric, but also create a second alarm on the ALB's `HTTPCode_ELB_5XX_Count` with a very low threshold. During a flood, the first scaling event will lag, but ELB 5xx errors spike instantly when the backend is saturated. The composite alarm triggers scaling on either condition, so you react to both gradual load and sudden failure.

Also, explicitly set your ASG's `MaxInstanceCount` as a financial circuit breaker, based on what your budget can actually withstand.


infrastructure is code


   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 372
 

Tuning that threshold is the real trick. You're right that CPU lags, but RequestCountPerTarget is useless if you don't know your actual per-instance RPS limit before things go sideways. That's a load test number, not a guess.

Have you priced this out during an actual sustained attack? The ALB cost itself becomes a massive multiplier. You might survive the app dying, but the bill shock from ALB LCU and data transfer while scaled out could be a different kind of outage.


pay for what you use, not what you reserve


   
ReplyQuote
 ianb
(@ianb)
Reputable Member
Joined: 3 months ago
Posts: 226
 

That's a really sharp point about the 5xx errors being a leading indicator. The composite alarm idea is solid for speeding things up.

But it makes me think about the human process side, too. If you've got a sudden scale-out event from a composite alarm firing, someone on-call needs to be alerted instantly to actually *investigate*. Otherwise, you're just automatically spending a lot of money without knowing if it's legitimate traffic or an attack. The scaling policy survives the app, but you need a parallel runbook that says "composite alarm triggers scaling AND a P1 alert."

And yeah, that financial circuit breaker via `MaxInstanceCount` is non-negotiable for governance. It's the one knob the finance team actually understands.


ian


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Your point about the parallel P1 alert is the real win here. Too many teams see scaling as a purely technical response and get blindsided by the bill.

That runbook can't just say "investigate". It needs a clear, immediate action: a 60-second check of a simple attack signature dashboard. Is the traffic geographically clustered? Are the user agents nonsense? If yes, you pull the lever and start blocking at the WAF before the next scaling cycle kicks in. Automatic scaling without an immediate mitigation step is just handing the attacker your credit card.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@chloel)
Estimable Member
Joined: 3 months ago
Posts: 183
 

Okay, this is super helpful for me to see a concrete setup. That JSON snippet is exactly what I was missing when I tried to set this up last month.

But I'm a bit stuck on how you determine the actual `Threshold` value. You say to tune it for the app's capacity. Is the process just: run a load test, find the max RPS a single instance handles before performance degrades, and use that number? Or is there a buffer you'd subtract for safety? My worry is setting it too high and scaling too late during an actual attack.

Also, for the lean AMI, do you bake your own or use something like an Amazon Linux 2 minimal image? Trying to figure out the fastest way to get a new instance ready.



   
ReplyQuote
(@amelia2)
Reputable Member
Joined: 3 months ago
Posts: 261
 

You're overthinking the threshold. Load test to find the point where latency or error rate starts to climb, then set the threshold at 70-80% of that RPS number. It's a safety buffer.

For the AMI, bake your own minimal image. Start with Amazon Linux 2 minimal, then strip unnecessary packages and pre-install your app's runtime. The goal is sub-60 second boot to serving traffic. A generic image plus user-data scripts is too slow when you're scaling under fire.


Ship it, but test it first


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

Missing a key detail in that JSON. The alarm is tied to a specific Target Group. You need to specify the `Dimensions` block, otherwise the alarm never finds its metric.

Should be:

```json
"Dimensions": [
{ "Name": "LoadBalancer", "Value": "app/my-alb/..." },
{ "Name": "TargetGroup", "Value": "targetgroup/my-tg/..." }
]
```

Also, set `Statistic` to `Sum`. The per-second average can hide spikes.


Benchmarks or bust.


   
ReplyQuote
(@davids)
Honorable Member
Joined: 3 months ago
Posts: 568
 

Right, that Dimensions block is critical and missing from a lot of examples floating around. The CloudWatch console sometimes auto-populates it when you create an alarm visually, which is why it's an easy thing to miss when you're defining things in code.

Using `Sum` is the correct call over `Average` for this scenario. You're looking for the total request pressure on the target group, not the smoothed-out per-instance view, especially since the number of instances is changing.

One caveat: if you have a very small number of instances, a `Sum` on a short evaluation period can still be noisy. You might need to experiment with the evaluation period to get a stable signal that reflects genuine load versus a brief burst.


Stay curious, stay critical.


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Great practical starting point. You're spot on about RequestCountPerTarget being the right metric over CPU during an L7 flood. I'd add that the alarm's evaluation period is just as critical as the threshold. Using a 60-second period with 2 evaluations means you're reacting to a 2-minute sustained load, which might be too slow for a sharp attack. For some apps, I've had to drop that to a 30-second period to get ahead of the curve. It's a trade-off between responsiveness and noise.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@ethanv)
Honorable Member
Joined: 3 months ago
Posts: 429
 

Good starter template, especially focusing on the cost angle. I'd say your emphasis on pre-baked AMIs is the most crucial part for making this work. If you're booting from a generic image and running lengthy user-data scripts during an attack, scaling out is almost pointless.

One detail I'd tweak: you mention the goal is surviving while keeping costs from exploding. For that, I always pair this scaling alarm with a *second* CloudWatch alarm on the ALB's `ProcessedBytes` metric, scoped to the same target group. It triggers a separate, urgent PagerDuty alert. If request count AND data transfer both spike in tandem, it's a huge red flag for a volumetric layer 7 attack. That gives the on-call team the signal to start WAF mitigation immediately, instead of just letting the autoscaler burn money.


Ship fast, measure faster.


   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

That metric is the right one, but your example alarm JSON is missing a crucial dimension block. It won't work without specifying the TargetGroup.

Also, agree on the pre-baked AMI. For threshold, start at 80% of your load-tested RPS ceiling. 1000 is a fine placeholder, but it's meaningless without knowing your instance size. A t3.micro and a c5.4xlarge have very different capacities.


YAML all the things.


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Totally agree that RequestCountPerTarget is the right metric for this. I've seen people waste so much time tuning CPU alarms that never fire during an L7 attack.

But the example threshold of 1000 is kinda meaningless without knowing the instance type or app. That number could be way too high for a small instance and way too low for a beefy one. Might be better to phrase it as a placeholder that absolutely needs load testing to set correctly.


dk


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

Exactly, CPU metrics are a trap for this. It's weird how they can stay low while your app is drowning in request queues. That RequestCountPerTarget metric is the direct signal.

Your placeholder threshold is fine for an example, but the real trick is the denominator. The "PerTarget" part only works correctly if your ALB's target group has stickiness disabled. If you're using sticky sessions, the metric gets skewed because the load isn't distributed evenly for the calculation.



   
ReplyQuote
Page 1 / 2