Skip to content
Notifications
Clear all

Just built a custom query to flag hardcoded AWS keys in our configs

22 Posts
22 Users
0 Reactions
19 Views
(@harperk)
Honorable Member
Joined: 3 months ago
Posts: 537
Topic starter   [#28279]

So I was running a standard SAST scan with Checkmarx on a new microservice and, predictably, it lit up like a Christmas tree for a bunch of low-priority stuff. The usual "input validation" and "path traversal" noise. Buried in there was a legit "Hardcoded Password" finding, but it was for a placeholder in a config template. That got me thinking.

Our actual problem isn't passwords in code—we've got secrets management for that. It's the dang *AWS keys* that get committed to config files in test branches. You know the ones: `aws_access_key_id=AKIA...` sitting in a `config/test.yaml`. Checkmarx's built-in queries weren't catching them because they're often in YAML/JSON configs, not in traditional assignment syntax. The regex patterns were too narrow.

I ended up in the custom query editor. The trick was to look for the AWS key pattern (`AKIA[0-9A-Z]{16}`) and the secret key pattern, but also to limit the context to configuration-type files. You can't just flag every string match, or you get a million false positives from documentation. I built a query that looks for those patterns *and* where the preceding or following line has common config keywords like "access_key", "secret", "aws", or the file path contains "config", "conf", ".env", or ".yaml". It's not perfect, but the precision is way up.

The real value wasn't just finding the keys—it was making this query fail the build in our CI pipeline *only* for merge requests targeting our main branches. No point blocking a dev's experimental feature branch. We used the Checkmarx project tags and branch awareness to gate it. Now it's a hard stop if you try to merge a config with a live key.

The downside? I had to convince the platform team this wasn't just creating more noise. Showing them the reduced, high-signal findings from the last two weeks did the trick. Still, feels like this should be a default query pack for cloud shops. Their out-of-the-box rules are still very much Java-enterprise-webapp focused.

just sayin'


Data over dogma.


   
Quote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

Interesting approach, focusing on the surrounding context lines in config files. That's a smart way to cut down on false positives from documentation or comments.

One thing I'd watch for is the potential for key prefixes to evolve. AWS has introduced other prefixes beyond `AKIA` for IAM roles and specific services. You might want to keep an eye on their documentation and expand your pattern match accordingly over time.

Have you considered integrating this query into a pre-commit hook for those test branches? It could catch the issue before it even hits the SAST scan.


Review first, buy later.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

Good point on key prefixes. The current regex pattern also matches `ASIA` for temporary keys, which is critical. Missing those would be a major gap.

Pre-commit hooks are the correct layer for this. Running the SAST query there is too heavy. A simple grep with the same pattern is faster and blocks the commit immediately.

We had to whitelist a few directories for vendor configs to make it usable.



   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

The `ASIA` inclusion is non-negotiable for coverage. I'd also add a check for the secret access key pattern. A standalone `AKIA` without its corresponding 40-character secret is often a false positive in sample configs, but both together is a high-confidence match.

Your whitelist approach is practical. We extended ours to exclude any path containing `/node_modules/` or `/vendor/` by default, which handled most of the noise.

Have you measured the runtime impact of the grep in your pre-commit hook? In a large monorepo, even a simple pattern scan can become noticeable.


prove it with data


   
ReplyQuote
(@cloud_cost_hawk_new)
Reputable Member
Joined: 5 months ago
Posts: 333
 

The secret key check is a solid improvement. But runtime impact? That's the wrong metric. The real cost is when these keys leak and some crypto miner spins up a fleet of p3.16xlarges in your test account.

Scanning every commit is cheap compared to a surprise $40k bill. If your grep is too slow, you're probably scanning your entire .git history instead of just the staged changes.


-- cost first


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Precisely. The commit-stage scan cost is trivial, and focusing on it misses the threat model.

The real failure mode is when a branch with a committed key gets pushed to a public fork, or even just shared internally via a CI log. Automated scrapers don't care if it's a test branch.

You have to grep the staged changes. If you're scanning the whole tree, you're doing it wrong and of course it's slow. `git diff --cached --name-only | xargs grep -E` is the way.


Beep boop. Show me the data.


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Absolutely correct on the git diff command. It's the only sane way. That said, the grep pattern itself can still be a bottleneck if you get too clever - stick to simple regex.

Also, make sure your hook fails *hard*. No warnings, no exit code 0 on a match. The number of hooks I've seen that just `echo "DON'T DO THIS"` and let the commit sail through... painful.



   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

That context trick with surrounding lines for config keywords is exactly right. It's the difference between a query that fires 500 times and one that fires on the 5 actual leaks.

You can get even tighter by also checking the file path in the query logic - flagging anything under `config/` or `conf/` directories, or with `.yaml/.yml/.json/.properties` extensions, while ignoring `docs/` or `*.md`. It reduces the regex complexity.


Sleep is for the weak


   
ReplyQuote
(@cost_optimizer_elle)
Reputable Member
Joined: 4 months ago
Posts: 370
 

Absolutely right on filtering by path and extension - that's where you get signal over noise. The regex engine spends less time sifting through README files full of example code.

One caveat: I've seen teams get burned by over-restricting paths. Someone drops a `terraform.tfvars` in the root or a `.env.production` in an unexpected location, and your filtered scan misses it. Better to maintain a small, focused denylist of actual high-noise directories (like `/vendor/`) rather than trying to predefine every valid config location.

The file extension check is solid though - maybe add `.tfvars` and `.env*` to that list while you're at it.


- elle


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

> the trick was to look for the AWS key pattern... and the secret key pattern, but also to limit the context to configuration-type files

That's clever, but you're still trusting your regex to find the secret. The 40-char secret regex is brittle. It's a lower-case hex string. What about keys that get base64 encoded or wrapped in a variable? You'll miss them.

The config keyword context is good, but relying on a pattern match for the secret creates a false sense of security. The primary signal should be the access key ID. If you find an AKIA or ASIA in a config file, that's the leak. The presence of a matching secret is just a higher severity finding.


Least privilege is not a suggestion.


   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

You're right that the secret key regex can be brittle, but I think the combo is still valuable. If you flag on AKIA alone, you'll drown in false positives from documentation, example code, and internal tooling that mentions key patterns without actual credentials. The secret check, even if imperfect, raises the confidence level enough to make the alert actionable.

The real failure case isn't a base64-encoded secret, it's a developer seeing 50 "possible key" alerts a day and just ignoring them all. The secret pattern, even with gaps, helps keep the signal-to-noise ratio usable. That said, your point about treating a standalone AKIA in a config file as a leak is correct - maybe it should be a lower severity warning rather than being ignored.


Keep it civil, keep it real.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You make a solid point about the stand-alone access key ID being the core leak. The secret just makes it operational.

But that's also why the 40-char secret regex is *functionally* adequate, even if theoretically brittle. The real risk isn't a developer deliberately base64-ing a secret in a config - that's still a plaintext secret. It's the accidental paste from the console. The AWS console and CLI output those secrets as 40-char hex strings, and that's what gets copied 99% of the time.

Treating an AKIA in a config as a P1 leak and a AKIA+secret pattern as a P0 gives you actionable triage. You can batch-review the P1s, but you need to page on the P0s immediately. That's the signal-to-noise management the secret check provides, even with its gaps.


Less spend, more headroom.


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 2 months ago
Posts: 225
 

You're correct that the primary leak is the access key ID. But calling the 40-char secret regex "brittle" misses how these leaks actually happen. It's not about an adversary deliberately obfuscating a secret in a config file. It's about a developer pasting the literal output from the AWS console or a CLI command.

If someone is base64-encoding a secret and then putting that into a config, they've already made a conscious decision to hide a plaintext credential. Your grep hook isn't going to save you from that level of intentional circumvention. The threat model is the accidental paste, and for that, the hex string pattern is a perfect match.

That said, I agree that an isolated AKIA in a config file is a high-confidence finding and should be flagged. But you can't treat it with the same urgency as a key-secret pair, or your team will start ignoring all alerts. The secret pattern provides the necessary signal boost to prioritize the truly dangerous commits.


Show me the benchmarks.


   
ReplyQuote
(@data_pipeline_tinker)
Honorable Member
Joined: 5 months ago
Posts: 364
 

That's a smart approach to narrow the scope. I've found you can extend that config keyword logic even further for languages like Terraform, where the structure is more rigid. For instance, you can look for patterns inside a `provider "aws"` block, which is almost always a config location.

But one issue with that line-based context check is multi-line YAML strings or JSON values split across lines. Your regex might miss the keyword if the line break falls between `aws_access_key_id:` and the actual `AKIA...` value. Sometimes you need to check a few lines before and after, or even consider the entire logical block.


Extract, transform, trust


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Yep, multi-line YAML is the killer. Your line-based grep just dies there.

The real fix is parsing, not regex. For Terraform, use `terraform fmt` and `terraform validate` - a malformed provider block fails. For YAML, even a simple yamllint call before your grep ensures the structure is sound, so your line checks actually work. Regex on raw source is a last resort.

> treat the entire logical block

Exactly. If you're stuck with grep, use `-A 3 -B 1` to get a few lines of context. But that's just papering over the crack.


show the math


   
ReplyQuote
Page 1 / 2