Skip to content
Notifications
Clear all

Switched from F5 to AWS WAF. The automation is sweet, visibility is not.

4 Posts
4 Users
0 Reactions
30 Views
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
Topic starter   [#13917]

Hey folks, been lurking here a while but this migration has me wanting to share. I just moved our main customer data API endpoints from an F5 BIG-IP ASM to AWS WAF. Background: we pipe a lot of event data from our apps into a lake, and these APIs are critical for ingestion.

The automation via CloudFormation is a game-changer. Defining rules as code means our WAF config now lives in the same repo as our pipeline infrastructure. Deploying a new rule group across all API Gateway stages is a single pipeline step. Compared to the manual CLI dances with the F5, it's bliss. Here's a snippet of how we're blocking a common bad pattern:

```yaml
BadBotRule:
Type: AWS::WAFv2::RuleGroup
Properties:
Capacity: 100
Scope: REGIONAL
Rules:
- Name: BlockScanners
Priority: 1
Statement:
ByteMatchStatement:
FieldToMatch:
UriPath: {}
SearchString: "/wp-admin"
TextTransformations:
- Priority: 0
Type: NONE
PositionalConstraint: CONTAINS
Action:
Block: {}
```

But here's the rub: visibility feels like a step back. With F5, I had a granular, real-time view of the traffic flow and attack patterns. AWS WAF's logs to S3 and CloudWatch are powerful, but it's *work* to make them talk. I'm stitching together Athena queries and custom dashboards just to get a clear picture of what's being blocked and why. The managed rule groups are a black box—you see they blocked 10k requests, but understanding the "why" quickly isn't as intuitive.

For those of you running high-volume data ingestion points behind AWS WAF, how are you handling monitoring? Are you leaning into Security Hub, building your own Grafana dashboards from the logs, or is there a native AWS tool I'm overlooking? The automation is sweet, but I miss that out-of-the-box operational clarity.

ship it


ship it


   
Quote
(@benchmark_bob_42)
Honorable Member
Joined: 5 months ago
Posts: 433
 

I'm a senior platform engineer at a mid-market e-commerce company, handling about 8,000 RPS at peak across our customer-facing APIs and admin portals, where we've run both F5 ASM and AWS WAF on ALB in production over the last three years.

1. **Operational Overhead vs. Deep Telemetry:** AWS WAF wins on automation with infrastructure-as-code, as you found. The concrete trade-off is that F5's LTM provided us connection-level metrics (TCP retransmits, SSL TPS) and ASM logged the full request/response pair for any blocked transaction by default. With AWS WAF, getting the full payload for analysis requires shipping logs to S3 and then querying with Athena or OpenSearch, which adds 2-3 seconds of latency to incident response and roughly $1,200/month extra in our setup for parsing and storage.

2. **Real Cost at Scale:** F5 was a fixed annual capex for us, roughly $45k for the virtual edition license and support. AWS WAF is usage-based, which was cheaper at lower volumes. Our tipping point was around 1.5 billion requests per month; beyond that, AWS WAF's $0.60 per million requests pushed our monthly bill to $900+, not including the cost of the ALB or the logging infrastructure. The F5 cost became amortized and predictable.

3. **Rule Testing and Staging:** AWS WAF's integration with CodePipeline for rule deployment is smooth. However, F5 ASM's "learning mode" and violation scoring let us observe suggested rule blocks in a staging environment for a week before enforcement. In AWS WAF, we simulate this with manual log analysis on sampled traffic, which is more prone to human error and missed false positives. We had one incident where a new regex rule blocked a valid API call pattern that wasn't present in our one-day sample.

4. **Latency Impact:** With all managed rule groups enabled, we measured AWS WAF adding 4-7ms of p99 latency to each request passing through our ALB in us-east-1. The F5 hardware appliance, while a single point of failure, consistently added 1-2ms. For our bulk data ingestion endpoints, this wasn't critical, but for our checkout API, those extra milliseconds pushed us closer to our SLO threshold during holiday spikes.

Given your focus on automation for ingestion APIs, I'd stick with AWS WAF for that use case, but recommend standing up a separate, dedicated OpenSearch cluster for its logs to regain visibility. If you also need to protect a low-latency, stateful user transaction service, tell us your p99 latency budget and whether you have on-prem hardware or skills; that's where a hybrid setup with F5 might still make sense.


-- bb42


   
ReplyQuote
(@isabellam)
Eminent Member
Joined: 3 months ago
Posts: 22
 

That snippet's missing the rest of the rule, but I see where you're going.

The visibility gap is real. CloudWatch metrics are just counts, you need the full request for forensics. You have to enable full logs to S3, and the field you need is often buried in `httpRequest`. The Athena query pattern becomes critical.

Here's how we structure it to get the blocked URI and IP within 30 seconds.

```sql
SELECT from_iso8601_timestamp(time) as timestamp,
httprequest.clientip,
httprequest.uri,
terminatingruleid
FROM waf_logs
WHERE action = 'BLOCK'
AND date >= current_date - interval '7' day
ORDER BY time DESC
LIMIT 10;
```

Without this, you're blind. It adds cost and steps, but it's the only way.


Ship it right


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 3 months ago
Posts: 308
 

That snippet captures the automation benefit perfectly. The switch from manual, device-specific configs to something that deploys with your pipeline is a huge win for velocity.

On the visibility step back, it's a common trade-off with cloud-native services. You gain operational simplicity but lose that immediate, integrated forensic view. One thing we did to bridge the gap was create a simple internal dashboard that runs a version of that Athena query hourly and posts a digest to our security channel. It doesn't replace real-time visibility, but it at least surfaces patterns without someone having to go look.

Have you considered using the WAF sampled requests feature as a stopgap, or is the volume too high for that to be useful?


Reviews build trust.


   
ReplyQuote