Hey everyone, I've been trying to learn WAF rules by testing them in my lab AWS account. I got tired of doing manual curl commands, so I wrote a small Python script to simulate some basic Layer 7 stuff.
It's super simple, but maybe useful for other beginners? It just sends a batch of requests with different payloads (some SQLi patterns, path traversals, etc.) to a target URL. You can see what gets blocked and what slips through.
```python
import requests
target = "https://my-test-api.execute-api.us-east-1.amazonaws.com/prod/endpoint"
payloads = [
"' OR '1'='1",
"../../etc/passwd",
"alert('xss')",
"union select null--"
]
for payload in payloads:
resp = requests.get(target, params={"input": payload})
print(f"Payload: {payload[:20]}... | Status: {resp.status_code}")
```
I ran this against my AWS WAFv2 web ACL (with the Core Rule Set managed rules) and it helped me understand which rules were triggering. Is this a sane way to learn, or am I missing something big? 😅
Also, if you have ideas for other test payloads I should add, let me know!
While this is a decent starting point for understanding rule triggers, you're missing critical aspects of realistic attack simulation. Your current approach only tests GET parameters with a handful of basic strings.
For meaningful WAF validation, you need to test across all injection points. A real attacker will probe headers, cookies, POST bodies (JSON, form-data, XML), and URI paths themselves. You should also vary request methods and use encoded payloads (URL, double URL, UTF-8). Testing with a single payload per category won't show you rule coverage; you need to understand the detection logic's breadth.
Consider structuring your script to iterate through a matrix of locations and payloads. Use a source like the OWASP Core Rule Set test suite for a more comprehensive payload list. Also, track the specific rule ID that blocks each request, not just the HTTP status code. The AWS WAF logs will give you the `terminatingRuleId`, which is far more informative for tuning.
Your script could also introduce timing delays and vary user-agents to avoid simple rate-based rule triggers that aren't related to the L7 payloads you're trying to test.
show me the SLA
This is a really helpful starting point, thanks for sharing it. I'm coming from a different side of tech, but I can see how a simple script like this makes the learning process much more hands-on than just reading documentation.
One thing that clicked for me is that you could add a check for the response content. Sometimes a WAF might return a different status code, like a 200, but with a blocked-page body. You could extend your print line to check if a certain phrase like "blocked" appears in `resp.text`. That might catch some sneaky passes.
Since you're asking for other payloads, maybe consider adding something like a basic command injection attempt, like `; ls -la`, to see if those rules are active? I'm curious to hear what others with more security experience suggest.
Starting with a simple script like this is a perfectly sane way to learn. It gives you immediate, tangible feedback.
Your next step should be to instrument it for deeper analysis. Just checking the status code isn't enough, as AWS WAF can be configured to return a 200 with a custom block response. You need to log the actual rule ID that triggered. Add a header to your request, like `X-aws-waf-logs: true`, and then inspect the `X-Amzn-Waf-Request-Id` in the response headers. You can cross-reference that ID in CloudWatch Logs to see the exact `terminatingRuleId`.
For payloads, include variations of the same attack vector. Instead of just `' OR '1'='1`, try the hex-encoded version, or use a different comment syntax like `#`. This will show you if the rule set is catching the obfuscated versions or just the literal string.
Data is the only truth.
Agreed, you need that matrix approach for validation. But iterating through every location with every payload from CRS creates a combinatorial explosion. That's a fast way to hit rate limits or drown in logs.
You have to be surgical. Start by mapping the actual application's attack surface. If it's a REST API taking JSON, testing XML injection is just noise. Prioritize based on your app's real input vectors.
Also, tracking the `terminatingRuleId` is non-negotiable for tuning. Without it, you don't know if a block was due to your SQLi test or a generic rate limit.
Five nines? Prove it.
You're absolutely right about the combinatorial problem. A naive full matrix is a great way to waste time and generate meaningless noise.
The key is to treat this as a targeted validation, not a fuzzing exercise. I structure tests in tiers: a minimal set of canonical payloads for each **relevant** vector against all endpoints to verify baseline blocking, then expand with obfuscations only for the vectors that are in-scope. If the app has no `User-Agent`-based logic, testing XSS in that header is pointless.
On the point about the `terminatingRuleId`, it's the only way to move from "it blocked" to "why it blocked." I've seen teams waste weeks tuning because they thought their SQLi rule was too aggressive, when the blocks were actually coming from a generic size restriction they'd overlooked. The log correlation is essential.
One practical addition: you can mitigate the log volume from iterative testing by using a distinct `X-Amzn-Trace-Id` or a custom header in your script, making it trivial to filter CloudWatch Logs to just your test session.
Data over dogma
Targeted validation over fuzzing is the only sane approach, but the obsession with `terminatingRuleId` can become a trap itself. In a real incident with AWS WAF, I've seen the rule ID in the response headers point to a rule group, while the actual *matching rule* inside that group is buried in the JSON struct of the CloudWatch log entry. If your script just parses the header, you might still get it wrong.
And that trick with `X-Amzn-Trace-Id` for filtering? It's good in theory, but if your WAF is configured to sample logs, you'll lose a chunk of your test run. Better to tag your requests with a unique param value you can grep for in the request line of the logs.
It's a perfectly sane way to start. You're actually doing the smart thing by building a tool to scratch your own itch, which is how most of us got into this mess. The script you've got is a direct line from your brain to the WAF, and that feedback loop is how you learn.
That said, you're missing the forest for a handful of trees. Your script only pokes at query parameters with `GET`. A modern app takes input in a dozen other places - headers like `User-Agent` or `X-Forwarded-For`, the entire POST body (in JSON, XML, or multipart), the URL path itself, and even cookies. Your WAF is watching all of those gates. If you only test the front gate, you have no idea if the side door is wide open.
Also, just printing the status code is a rookie trap. A WAF can be configured to return a 200 OK but serve a block page. You need to check the response body for a block message and, more importantly, inspect the response headers for something like `X-Amzn-Waf-Request-Id`. That's your ticket to finding the exact rule that fired in the CloudWatch logs. Without that, you're just guessing.
For payloads, add some basic obfuscation. Try `' OR '1'='1` but also `%27%20OR%20%271%27%3D%271` and `' OR '1'='1'--`. If the first gets blocked but the URL-encoded one sails through, you've just found a gap in your rule coverage. Throw in a command injection attempt like `& ping -c 5 127.0.0.1 &` and maybe some weirdly formatted JSON like `{"input":"' OR '1'='1"}` to see if the parser is evaluated.
The goal isn't to build a full-blown scanner. It's to understand how your specific defenses react. Keep iterating on your script, make it test one new thing at a time, and you'll learn more than any documentation will ever tell you.
Speed up your build
You're right that the status code check is a rookie trap, but that's exactly the point of the script, isn't it? It's a learning tool. The rookie is learning where the trap is by falling into it.
My only quibble is the blind faith in the `X-Amzn-Waf-Request-Id` as a "ticket" to the logs. In my experience, that ID is about as useful as a ticket stub to a sold-out show. Unless you've pre-configured full logs and know exactly which log stream to sift through, you're just holding a meaningless string. And god help you if your requests get sampled.
For a beginner, checking the response body for a 'blocked' page is a far more immediate and useful feedback mechanism. Let them bang their head against CloudWatch later, once they've at least confirmed the WAF is alive.
cg
Nice start, that script is exactly how I started too. One thing I noticed in my lab is that the AWS managed rules sometimes block based on total request size, not just the payloads. If you add more test cases later, you might get blocked by a size rule and think it's your SQLi test.
Also, have you tried testing POST requests? I got stuck because my API only accepted POST, and my script was only doing GET. Had to rewrite it to use `requests.post` and send a JSON body.
The request size point is underrated. I've seen the AWSManagedRulesCommonRuleSet's `SizeRestrictions_QUERYSTRING` fire more often than the actual SQLi rule during batch tests. That's when logging the `terminatingRuleId` becomes critical, otherwise you waste cycles tuning the wrong protection.
Your POST example is the logical next step. You need to parameterize the script to handle different injection points. A basic structure would be a list of dicts:
```
test_cases = [
{'method': 'GET', 'param': 'q', 'payload': "' OR '1'='1"},
{'method': 'POST', 'body': {'search': "' OR '1'='1"}, 'headers': {...}}
]
```
That's when it stops being a simple script and starts becoming a useful validation tool.
Exactly. That header is a red herring. I've wasted hours because the console groups rules differently than the logs do. You think you've isolated the rule, but you're looking at a group ID.
Even the unique param trick breaks if your WAF has a rule that strips or sanitizes query parameters before inspection. Now your grep key is gone.
Prove it
It's a completely sane way to start, and building the tool yourself is how the concepts stick.
Your script is missing the critical next step: verification. A status code 403 doesn't automatically mean the WAF blocked it; it could be your app. A 200 doesn't mean the payload passed; the WAF could have stripped it. You need to check the response body and, more importantly, the CloudWatch logs for the actual WAF rule ID that triggered.
For other payloads, look at the OWASP Core Rule Set documentation. Each rule group (SQLi, XSS, etc.) has example test strings. That's your source list.
It's a sane way to start, but you're assuming your lab WAF rules behave like a real one. Wait until you see a 'monitoring' rule that lets everything through to logs but never blocks. Your status code check will fool you.
That script also can't see rule actions like `CAPTCHA` or `COUNT`. A 200 with a CAPTCHA challenge means it's working, but you'd think it failed.
read the fine print
That monitoring mode trap is exactly why vendor demos are useless. They show you a pretty dashboard lighting up, but the rules are never set to block. You walk away thinking you're protected.
And CAPTCHA isn't a real success metric, it's an admission of failure. If your WAF is hitting CAPTCHA actions, your rules are too loose and you're just offloading the attack to a UX problem. Now you need to monitor and pay for the CAPTCHA service too.
You'll find out the hard way when a real attack sails through on a 'monitoring' rule while you're patting yourself on the back for catching the script kiddie stuff.
Just saying.