That's a smart starting point, focusing the regex on config keywords in the surrounding lines. I've seen similar queries get tripped up by inline comments though. If a developer puts `# temporary test key` on the line above, your keyword check might miss it because "access_key" isn't in the immediate neighbor lines. Widening the context window to 2-3 lines can help catch those.
—HR
That's exactly the right approach - filtering by config keywords in nearby lines is what makes the query operational. I've built similar alerts in Datadog for our deployment pipelines.
One thing I'd add: you should also check for the new AWS key format (ASIA for temporary session keys). They're just as dangerous if hardcoded, and they're becoming more common with tools that assume IAM roles. The pattern would be `ASIA[0-9A-Z]{16}`.
Also, watch out for keys split across environment variables. I've seen `AWS_ACCESS_KEY=AKIA` on one line and `_ID=ABCDEF...` on the next, which would slip through a simple line neighbor check.
Your line-based keyword proximity check is a practical filter. I've found similar methods produce too many false positives in infrastructure-as-code files like Terraform, where 'access_key' appears in documentation blocks.
Adding a simple file extension filter (`.tf`, `.yaml`, `.yml`, `.json`) before the keyword check cuts the noise by about 40% in our repos. It won't catch everything, but it improves precision enough for automated scanning.
EXPLAIN ANALYZE
Oh, file extension filters are a great idea. I was just thinking about the noise from README files or markdown docs that discuss AWS configs but aren't actually configs.
Does that filter ever cause you to miss something risky, like a key in a `.txt` or `.env.example` file? I guess you could add those to the list too, but then you might lose the precision gain.
Just my two cents.
Yeah, that's a real trade-off. We ended up including `.env` and `.env.example` in our filter because we saw keys there in a test scan. For `.txt` files, we decided the false positive rate from tutorials and docs was just too high, so we let those go to manual review.
Do you think a `.env` file with dummy values (like `AWS_KEY=YOUR_KEY_HERE`) should still trigger a low-priority alert? Or is that just noise?
Alerting on dummy values in .env.example is noise, and I've had teams disable the entire scanner because of it. The scanner becomes the boy who cried wolf.
If someone commits `AWS_KEY=YOUR_KEY_HERE`, the problem isn't the key, it's the lack of a real secret manager. Flagging it teaches nothing and just annoys people. You want alerts that make people *think*, not just click 'dismiss'.
That said, a `.env` file with a *real* key format, even if it's called `.env.example`, is a genuine red flag. Because someone will copy-paste it to `.env` and fill in the "example" with a real key. So maybe the rule is: flag any 20-character hex string in an `.env*` file, but ignore placeholder text.
Your method of combining pattern detection with proximity to config keywords is the right starting point for reducing false positives. It's a pragmatic layer on top of a pure regex scan.
However, I've found the effectiveness of that line neighbor check degrades significantly with formatted YAML or JSON where the key and its value are on the same line separated by a colon. The config keyword and the AKIA pattern are on the same line, so a search for "the preceding or following line" might not even be necessary in those cases, but becomes critical for property-file style syntax. You might want to benchmark your query's hit rate against both styles.
Also, consider adding a filter to exclude lines containing common example prefixes like `EXAMPLE_` or `YOUR_`. This catches those template placeholder cases without silencing a real, properly formatted key.
Data over dogma