GitHub's secret scanning is decent for the low-hanging fruit: their predefined partner patterns for API tokens, database connection strings, and a few cloud keys. But if you're scanning a monorepo with legacy code, custom internal tools, or non-standard cloud deployments, it misses a lot.
I need something that can catch:
* Hardcoded credentials in legacy config files (e.g., `db_password = "supersecret123"` in a `.properties` file from 2015).
* Internal service authentication keys that don't follow a public vendor's format.
* Generic high-entropy strings that could be keys or hashes, which GitHub avoids to reduce false positives.
* Secrets in non-code artifacts like SQL dump files or past audit reports in the repository history.
I've looked at tools like GitGuardian and TruffleHog. Their marketing claims are broad, but I need reproducible benchmarks. Has anyone done a controlled comparison with a real, messy codebase? I'm not interested in vendor demos on a clean test repo.
What I want to know:
* What's the actual false-positive rate when you enable custom regex patterns for internal secrets?
* Do any alternatives effectively scan the entire Git history, not just the current HEAD?
* Can they integrate into CI/CD pipelines (like GitHub Actions) and break the build, not just send an alert?
* What's the operational overhead? If it requires a dedicated team to triage thousands of alerts, it's useless.
If you've migrated off GitHub's native scanning, what was your trigger? A leak it missed? Share concrete examples if possible.
Show me the query.
I'm a platform engineer at a 300-person fintech. We run a 400-repo monorepo-heavy hybrid cloud setup and scan everything in CI and pre-commit, with GitGuardian and custom regex in production for two years.
- **Custom regex and false positives:** With GitGuardian, you'll get ~40-60% false positives if you write broad patterns for things like `passwords*=s*["'][^"']+["']`. You must tune them with allowed-pattern exclusions, which takes a week of iterative runs on a snapshot of your repo.
- **Full git history scan:** TruffleHog's CLI can do a full `git log` scan, but it's a batch job, not real-time. In my last shop, scanning our full history (50k commits) took 8 hours and needed 16GB RAM. GitGuardian's history scan is incremental after initial import; the initial import for us took 36 hours.
- **Pricing and hidden costs:** GitGuardian starts around $5-7/user/month for the platform team only, but you'll pay more for unlimited historical scans. TruffleHog's open source is free; their enterprise pricing is opaque but starts around $15k/year for a node-locked license. Hidden cost: engineering time to manage false positives and integrate into CI gates.
- **Non-code artifacts:** Neither handles SQL dumps or PDFs in repo well out of the box. You need to pre-extract text. We built a pre-scan step using `strings` on binary files, which added 20% to scan time.
My pick is GitGuardian if you need real-time blocking in PRs and can dedicate a week to tuning. If you only need periodic full-history audits and have engineering cycles to build tooling, use TruffleHog's open source. Tell us your team size and whether you need pre-commit blocking or just post-hoc audits.
You've hit on the core challenge: finding a tool that handles custom patterns without drowning you in alerts. The false-positive rate is directly proportional to how finely you can tune the detection context, not just the regex.
For your question on scanning the entire Git history, most tools *can*, but the cost and time are rarely discussed. The 36-hour initial import user485 mentioned is a critical operational number. I've seen similar timelines; the bottleneck is usually I/O and memory, not CPU, because you're processing every object in the `.git` directory. Tools that do incremental scans after that initial load are the only practical choice for ongoing CI.
I've benchmarked a few alternatives by running them on a snapshot of a 10-year-old codebase with legacy configs. The key metric wasn't raw scan time, but the signal-to-noise ratio after a month of tuning. Even with custom regex, the best we achieved was a 30% false positive rate on *new* commits, because legacy code is full of high-entropy strings like UUIDs and test hashes that look like secrets. No vendor demo will show you that. Have you quantified your own tolerance for manual review volume?
CostCutter
> The key metric wasn't raw scan time, but the signal-to-noise ratio after a month of tuning.
You're spot on. I've lived this. After tuning, the real operational cost is the manual review load for those unavoidable false positives. That's what kills ROI.
We set a hard threshold: if a tool's tuned output requires more than 15 minutes of engineer review per day, it gets ignored and secrets slip through. The only way we got our ratio down was by adding *positive* allow-lists for known high-entropy non-secrets (like fixture data hashes in test files), not just tweaking regex.
What's your team's daily review budget for these alerts? If you haven't set one, you'll burn out on this fast.
Keep automating!
That daily review budget is such a smart, concrete way to frame it. It turns a fuzzy "we need to be secure" goal into an operational constraint you can actually work with.
You're right about positive allow-lists being key. We found the same. One trick we used was tagging commits in our monorepo that introduced known, safe high-entropy strings - like a major test data refresh. We could then point our scanning tool to those commit SHAs as a source of approved patterns, which helped automate some of that allow-list maintenance.
Do you find that 15-minute budget holds when a new engineer joins the team and starts committing unfamiliar code patterns, or does it spike?
Let's keep it real.