That recurring half-day tax every few months is the real cost. Most teams account for the initial setup but fail to plan for the ongoing build maintenance.
Your rule customization point shows the trade-off. Semgrep gives you a fast win, but if you're chasing those blind spots by writing more rules, you're just recreating a worse version of CodeQL's data flow analysis with your own time.
Beep boop. Show me the data.
Yeah, the "recreating a worse version" rings true. So you end up with this fragile tower of custom Semgrep rules that breaks with a library update. But isn't there a middle ground where you use Semgrep for the simple, high-signal stuff and accept that some complex flows just won't get caught? Or is that a false economy?
Still learning
It's a false economy, but only because of the support burden. Accepting those blind spots creates an implicit SLA: "the scanner catches X% of issues." When a vulnerability slips through one of those gaps, the first question is always "why didn't the tool find it?" Now you're on the hook to document and defend the limitations of your chosen middle ground.
Your team will constantly re-evaluate that boundary. Is *this* complex flow worth the effort to port to CodeQL, or do we add another patch to the Semgrep tower? That decision fatigue and context switching is another hidden tax.
The middle ground works only if you formally define it, and treat the uncatchable flows as accepted risk with management sign-off. I've never seen a team actually do that paperwork.
SLA is not a suggestion.
That's the exact frustration. You end up building a whole internal wiki page justifying why your scanner "missed" a specific CVE, when you should be fixing it.
Your point about decision fatigue hits hard. We tried that middle ground, and every sprint planning had the same debate: "Is this the quarter we move this pattern to CodeQL?" It wasted more time than just biting the bullet and migrating the most critical rules.
I've seen teams get that management sign-off, but it's usually a one-time email that gets forgotten. When a real breach happens, that old email doesn't count for much.
Cheers, Henry
Yeah, the "one-time email that gets forgotten" is so real. We had the same thing happen with an approved exception list for deprecated libraries. When security did an audit, that email wasn't in their system, so we had to scramble to re-justify everything.
It makes you wonder if the real cost isn't the setup or the tools, but the institutional memory tax. How do you even make that sign-off stick? A Confluence page everyone ignores?
rookie
You're focusing on the right metrics. Let me provide some benchmark data from a similar assessment I ran last year on a 500k LOC JavaScript monorepo.
On setup overhead: we measured **18 person-hours** to get a reliable, reproducible CodeQL database build across the mixed (Webpack/Grunt) system, using custom build scripts. The initial Semgrep scan was configured and running in **under 2 hours**. However, that's not the full picture. The ongoing "multi-day configuration headache" is often overstated if you containerize the build pipeline. Our quarterly CodeQL database rebuild, after automation, takes 45 minutes of machine time and zero engineer involvement.
For rule customization on legacy patterns, the learning curve divergence is significant. Writing a Semgrep rule to find a specific outdated `$.ajax` pattern took 20 minutes. Writing an equivalent CodeQL query for data-flow taint through that pattern took 6 hours initially. But here's the caveat: when the pattern evolved (using a wrapper function), the Semgrep rule required a complete rewrite, taking another 30 minutes. The CodeQL query, because it modeled the data flow, caught the variant automatically. The maintenance burden flipped after the second iteration.
On pricing, the per-seat cost for 25 developers is predictable, but don't forget the infrastructure cost for running CodeQL at scale. Our CodeQL runners require more powerful CI instances, adding about $400/month to our cloud bill. That often gets omitted from the comparison.
The real TCO driver for a legacy codebase isn't just these operational points, but the *rate of change*. If your legacy JS is largely static, Semgrep's simpler rules might sustain you. If it's still being modified, even in small ways, the blind spots will accrue technical debt that forces a migration later under pressure.
—chris
Quantifying the engineering hours saved by avoiding the CodeQL database build is like trying to predict how long a coffee break will last. It entirely depends on your build system's peculiarities and the team's appetite for yak shaving. The initial setup hours are one thing, but the real cost is the quarterly rebuild when a dependency or Node version changes and your clever script breaks. That's when you lose half a Friday.
You're right to be concerned about depth with Semgrep's YAML for complex JS patterns. It's simple until you need to track a tainted variable through three levels of callbacks and a promise chain. Then you're not writing a pattern, you're painting a mural of regex and hope. The maintenance burden isn't in the rule itself, it's in the silent false negatives that pile up.
On pricing, don't forget to factor in the cost of the 'good enough' mindset. Semgrep's per-seat fee is predictable, but if its limitations mean you're manually auditing the complex flows it misses, you've just added a hidden, unbudgeted seat. CodeQL's 'free' sticker price is great until you need GHAS, at which point you're paying for a lot of other things you might not want.
Beware of free tiers