That's a really thorough test methodology, and focusing on the three dimensions you outlined is spot on. I'm especially interested in your findings on GCP rule coverage. We're heavily HubSpot and AWS, but we've started piloting some GCP projects for analytics workloads, and I'm already seeing the gaps you mentioned.
The compliance angle is critical for us too. Did you find the HIPAA and PCI-DSS rules in the community registry were practical for real-world Terraform, or were they more like generic placeholders that needed heavy customization? I worry about the "checkbox security" risk if the rules don't understand how compliance controls actually map to specific resource arguments in different clouds.
Also, curious on the integration dimension you mentioned - with GitHub Actions, was the config itself straightforward, or did you have to build a lot of wrapper logic to manage the findings output?
If it's not measurable, it's not marketing.
On the compliance rules, they were absolutely checkbox placeholders in our test. The PCI rule for GCP storage buckets flagged missing uniform bucket-level access, which is correct. But it was silent on the far more common misconfiguration we see: using the deprecated `acl` argument alongside `uniform_bucket_level_access = true`, which creates a conflict the rule didn't catch. So it gave a false sense of compliance.
For GitHub Actions, the basic config was trivial. The real work was in the wrapper logic to format the output and gate the pipeline. We had to write a custom script to fail only on high-confidence findings, otherwise the noise drowned the signal. Straight out-of-the-box, it fails on everything, which is what led to that 70% failure rate the earlier post mentioned.
Cloud costs are not destiny.
That's exactly the kind of rule that becomes a maintenance trap. You've essentially encoded your team's approval policy into Semgrep logic. What happens when the business changes that policy? Now you're not just updating a wiki, you're debugging a custom regex-like pattern to ensure the logic still maps correctly. You've traded a human review step for a brittle, version-controlled proxy that gives a false sense of automation.
It feels neat until you realize you're now in the business of maintaining a parallel, undocumented rule system that has to stay in perfect sync with actual human decisions. The cognitive load you saved the developer just got transferred to the rule author, permanently.
Skeptic by default
You've hit the nail on the head. We fell into that exact trap with a custom rule for VPC peering approvals. The rule was built around a specific business unit mapping that changed six months later. Updating the Semgrep pattern was a three hour debugging session because the logic had to match the new organizational structure, and we had to audit every existing PR that had previously passed.
The brittle proxy is real. It shifts the burden from a clear, documented approval workflow to a hidden technical debt. Now, a policy change requires a developer who understands both the business logic AND Semgrep's pattern syntax to untangle it. That's a much rarer skill set than someone who can just read and update a policy doc.
We've started treating custom rules like application code, with the same review and ownership requirements, because they encode business decisions. If you wouldn't put that logic in a Terraform module without tests, it doesn't belong in a Semgrep rule either.