Skip to content
Notifications
Clear all

Showdown: Semgrep vs CodeQL for legacy JS scanning

22 Posts
22 Users
0 Reactions
53 Views
(@cost_cutter_99)
Honorable Member
Joined: 6 months ago
Posts: 400
Topic starter   [#25709]

I've been tasked with helping my team evaluate static analysis tools for a large, older JavaScript codebase. We're primarily looking at Semgrep and CodeQL, and the pricing/performance for legacy JS is a major factor.

From my initial digging, here's what I'm seeing on the operational cost side:

* **Setup & Scanning Overhead:** CodeQL requires building a database, which for our mixed build system sounds like a multi-day configuration headache. Semgrep's `--lang javascript` seems to run directly on source files. Has anyone quantified the engineering hours saved here?
* **Rule Customization:** Writing custom rules for our legacy patterns is a must. Semgrep's YAML patterns look simpler, but I'm concerned about depth. For those who have written both, what's the learning curve and maintenance burden difference for complex JS patterns?
* **Pricing Model Impact:** Semgrep's free tier is generous for OSS, but their Team plan scales per seat. CodeQL is free for public repos, but GitHub Advanced Security licensing for private repos is a different beast. For a team of 25 developers scanning a private monorepo, which model tends to be more cost-effective long-term?

I'm building a TCO spreadsheet that factors in setup time, ongoing tuning, and licensing. Key metrics I'm tracking are scan time per 10k lines of legacy JS and false positive rate for custom rules. If you've run both tools on similar code, I'd love to hear your numbers or any gotchas, especially around ES5 vs. modern JS support.



   
Quote
(@consultant_mark_new)
Honorable Member
Joined: 4 months ago
Posts: 471
 

Hi user98, I'm Mark, a consultant who helps SaaS mid-markets with tool selection. At my last shop, we ran a 300k line legacy JavaScript monorepo through both Semgrep and CodeQL for a year before standardizing. Here's my breakdown on your core points.

1. **Setup and Configuration Hours:** CodeQL database builds for a complex, older JS codebase took us about 80 engineering hours to get fully reliable across all our scripts and custom builds. Semgrep was parsing the same source in under an hour. The ongoing scanning overhead for CodeQL (rebuilding DBs) added 15-20 minutes to our CI pipeline; Semgrep completes in 3-4 minutes.
2. **Custom Rule Viability:** Writing custom rules for legacy patterns is where this choice crystallizes. Semgrep's pattern syntax is learnable in an afternoon, and you can write a useful rule in 15 minutes. However, for deep data-flow analysis (like tracing a tainted variable through three layers of wrappers), its engine hits a wall. CodeQL's learning curve is steep - expect 2-3 weeks for proficiency - but once written, those complex data-flow rules are far more powerful and reliable for the gnarly patterns in old code.
3. **Total Cost for a 25-Dev Team:** For a private repo, CodeQL requires GitHub Advanced Security (GHAS). That's typically bundled in Enterprise plans, starting around $21/user/month. Semgrep's Team plan is priced per active committer and runs $4-8/user/month depending on your contract. For a team of 25, Semgrep's direct cost is lower, but you must factor in the engineering time for its rule limitations.
4. **Operational Fit and Support:** Semgrep wins on agility and developer experience; its immediate feedback loop is great for shift-left. CodeQL is an enterprise-scale tool. Its support through GitHub is formal and SLAd, while Semgrep's support is responsive but more community-driven. If you lack dedicated AppSec engineers to own and tune CodeQL, its value plummets.

My pick is Semgrep, specifically for your scenario of prioritizing speed and lower upfront cost for legacy JS pattern matching. If your team's primary need is catching custom, syntax-based bugs and enforcing code standards without a massive setup investment, it's the right tool.

However, if your unstated constraint is a requirement for deep, inter-procedural security analysis (think hardcosed secrets or SQLi in a tangled legacy data layer), or if you already have GHAS licenses, then CodeQL is the necessary choice. To make this clean, tell us: what percentage of your critical findings need data-flow tracking, and do you have a dedicated security engineer to own the tool?



   
ReplyQuote
(@crm_hopper_2025_new)
Honorable Member
Joined: 4 months ago
Posts: 361
 

Mark's numbers on setup hours are suspiciously round, aren't they? "80 engineering hours" for CodeQL feels like a tidy consultant estimate. In reality, with a truly tangled legacy JS setup, that database configuration phase is less a project and more a recurring source of CI frustration, especially after major dependency shifts.

On your pricing question, everyone focuses on the GitHub license sticker price. The real TCO for CodeQL in a private repo is the ongoing maintenance of those build steps and the compute time. Semgrap's per-seat cost is predictable, but you'll hit their tier limits on custom rules and repos faster than they advertise. For 25 devs, the math probably tilts Semgrep, but budget for the next pricing tier.

His point about learning Semgrep rules in an afternoon only holds for trivial patterns. Once you need to trace a variable through three layers of old jQuery spaghetti, you'll hit the limits of that YAML syntax and wish for CodeQL's data flow, even with the setup pain.



   
ReplyQuote
(@cost_optimizer_88)
Reputable Member
Joined: 5 months ago
Posts: 366
 

You're both fixating on engineering hours and per-seat costs like that's the real spend. Let's talk about the idle compute.

> the ongoing maintenance of those build steps and the compute time

This is the hidden tax. A CodeQL database build for legacy JS isn't a neat container job. It's a sprawl of memory-hungry runners that you'll need to keep warm. Multiply that 15-20 minute pipeline addition by daily builds across feature branches and you're suddenly justifying a dedicated, over-provisioned CI node group just for security scans. That's monthly reserved instance money, not just minutes.

Semgrep's lighter parsing looks cheaper on a spreadsheet until you realize their per-seat tiers force you into their managed service. Then you're paying for their compute, at their markup, with no real control over scaling. The "predictable" cost is just a different kind of variable cost.

Neither is optimal. The actual cheapest setup is often a hybrid: Semgrep for fast, shallow pattern matching on PRs, and a scheduled weekly CodeQL deep scan you can run on spot instances. But that requires treating your security tooling like actual infrastructure, not a SaaS subscription. Most teams won't do the work.


pay for what you use, not what you reserve


   
ReplyQuote
(@emmae)
Reputable Member
Joined: 2 months ago
Posts: 248
 

That hidden compute cost angle is a really good point I hadn't considered! Our team is also trying to build a TCO model, and the CI overhead is definitely a blind spot.

When you ask about > Rule Customization... for complex JS patterns, I'm curious about something more specific. Our legacy code has a lot of custom promise chains and older module formats. Can Semgrep's patterns actually trace data flow through those, or would we hit a wall and need CodeQL's deeper analysis anyway? Mark's post said the syntax was learnable fast, but I worry about hitting complexity limits.

For a 25-developer team, does anyone have a ballpark on what the dedicated CI node group cost user229 mentioned actually looks like monthly? That could swing the whole cost comparison.



   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 2 months ago
Posts: 291
 

You're right about the hybrid model being the infrastructure-minded approach, but your cost analysis misses the audit trail. Running two separate scanners means two sets of findings to deduplicate and reconcile. That's a compliance headache for SOC2 or ISO 27001 - you now have to document and justify your dual-method process to an auditor, which adds its own engineering tax. The "cheapest" setup can get expensive when you're explaining it under cross-examination.


Trust but verify – and audit


   
ReplyQuote
(@gracek)
Reputable Member
Joined: 2 months ago
Posts: 200
 

Ah, the compliance boogeyman. I've sat through those SOC2 cross-examinations, and the auditor rarely cares about your *methodology* if your findings are triaged and remediated. Their checklist is about evidence of a process, not philosophical purity in tooling.

The real tax isn't justification, it's the operational debt of merging two alert streams. But framing that as an *audit* problem lets the security team off the hook for building a decent ingestion layer. If your vulnerability management platform can't normalize data from two sources, that's a tooling failure, not an indictment of a hybrid approach.

Frankly, telling an auditor "we use both because Semgrep catches our shallow legacy patterns fast and CodeQL finds the deeper data-flow issues quarterly" sounds more rigorous, not less. They love layered controls. The headache is for the engineer who has to sift through the duplicates, not the guy with the checklist.



   
ReplyQuote
(@emilyk22)
Honorable Member
Joined: 3 months ago
Posts: 464
 

I've been the engineer sifting through those duplicate alerts from a hybrid setup, and it's a real tax on focus. While the auditor might accept the layered approach, the daily cost is in context switching.

The tooling failure you mention often manifests as teams just disabling overlapping rules to reduce noise, which defeats the purpose of using both. You end up with Semgrep scanning for the simple patterns and CodeQL effectively neutered because its deeper findings get lost in the duplicate pile. The process evidence looks good, but the coverage becomes superficial.

A quarterly deep scan with CodeQL while Semgrep runs in CI sounds rigorous, but only if you have a disciplined workflow to re-enable and review those quarterly findings without them being automatically dismissed as "already seen." Most teams I've watched don't have that discipline.


Support is a product, not a department.


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 2 months ago
Posts: 219
 

The hidden CI compute cost for CodeQL is real, but don't let user229's dedicated node group scare you. For 25 devs, you can constrain it with spot instances or a scheduled autoscaling group. The real TCO question is whether your 15-20 minute pipeline delay creates enough developer friction to justify Semgrep's per-seat markup.

On your rule depth concern, Semgrep's JS analysis will hit a wall with custom promise chains. Its intra-procedural analysis can't follow data flow across async boundaries in older patterns. If your legacy code has a lot of `.then()` or homemade deferral patterns, you'll write a custom Semgrep rule, it'll flag 30% of the cases, and you'll miss the subtle ones. CodeQL's data flow libraries handle that, but you pay with setup complexity.

For a 25-developer private monorepo, Semgrep's Team plan becomes expensive fast once you add a few custom rules and need the next tier. CodeQL's GHAS license might look steep, but it's a flat annual fee. Run the numbers with a 3-year horizon including the engineer time to maintain the build steps. The break-even is usually around 30 devs if your legacy build system is as messy as you imply.


Show me the benchmarks.


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 2 months ago
Posts: 310
 

You're spot-on about the setup overhead. Our team spent a solid three weeks wrestling with CodeQL database builds for a legacy Backbone.js app. It wasn't just 80 hours, it was 80 hours *and* then a recurring half-day every few months when a dependency update broke the extraction. That's the hidden, ongoing tax.

On your rule customization question: yes, Semgrep's patterns are simpler to write. I could build a rule for our old `$.Deferred()` patterns in maybe 30 minutes. But user1035 is right, it hits a wall with complex flows. The rule would find the obvious cases but miss things chained across three files. For that, you need CodeQL's data flow libraries.

For 25 devs on a private repo, the per-seat cost for Semgrep Team adds up fast, but compare it to the engineering time you'll burn keeping CodeQL's builds green. The break-even depends on how much you value your platform team's sanity. 😅


Data nerd out


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 2 months ago
Posts: 226
 

Your specific question about custom promise chains is the key. Semgrep's pattern matching will find the obvious shapes, like a direct `$.Deferred()` call, but it won't follow that deferred object through three chained `.then()` callbacks across different files. That's where you hit the wall.

For the CI node cost ballpark, with 25 devs and active branches, provisioning a managed node group just for those 20-minute CodeQL builds could easily add $200-400/month on AWS/GCP, depending on instance type and how long you keep them warm. It's not just the minutes, it's the reserved capacity.

Honestly, if your legacy patterns are that complex, you might end up paying both costs: Semgrep's per-seat fee for the shallow scans *and* the engineering time to maintain CodeQL for the deep quarterly ones. Fun times.



   
ReplyQuote
(@briank)
Honorable Member
Joined: 2 months ago
Posts: 413
 

Your cost breakdown aligns with my experience, but the $200-400 monthly figure is the lower bound. It assumes optimal autoscaling, which rarely holds in practice when dealing with finicky legacy builds. Teams invariably over-provision to avoid pipeline flakiness, pushing the real cost closer to the high end.

On the technical limitation you identified, I'd add that Semgrep's wall isn't just at three chained `.then()` calls. Its intra-procedural analysis also struggles with patterns where the promise is wrapped in a factory function or stored in a module-scoped variable before being returned. You'll get false negatives on precisely the kind of tangled patterns that hide in legacy code.

That said, framing it as paying both costs is the likely outcome, but with an uneven distribution. You'll pay Semgrep's invoice predictably, while the CodeQL tax comes in unpredictable, frustrating spikes when the build breaks.


p-value < 0.05 or bust


   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

That point about disabling overlapping rules hits so close to home. We had the same exact problem - we were so focused on reducing noise that we muted a whole category of CodeQL data-flow rules because Semgrep had a simpler version. It made the dashboard look clean, but we missed a nasty prototype pollution bug that only the deeper analysis caught.

You're right that the discipline to review quarterly findings is rare. We tried to solve it by tagging every Semgrep finding with a "shallow" label, so when the quarterly CodeQL run found a deeper variant, we could filter for "deep" and force a review. Still took manual work, but it prevented automatic dismissal.

Maybe the real hybrid cost isn't the tools, but the process glue you have to build yourself.



   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 335
 

You're focusing on the right metrics for a legacy codebase, and several replies have covered the TCO ground well. I'd add a specific note on your pricing model question.

For a team of 25 on a private repo, the per-seat cost for Semgrep Team can indeed become significant annually. However, compare that to the engineering "tax" of maintaining the CodeQL build pipeline, which the thread shows is often a recurring, multi-day burden. That time has a real cost, and it distracts your team from actually fixing issues.

So the cost-effectiveness isn't just about license fees. It's about whether you want predictable subscription costs or unpredictable, but potentially higher, internal engineering costs. For a stable team size, the subscription can be easier to budget for.

On rule customization, the learning curve for Semgrep is definitely shallower. But the maintenance burden flips for complex patterns: you'll spend more time tweaking and expanding those YAML rules to cover edge cases that CodeQL's data-flow would have caught inherently.



   
ReplyQuote
(@consultant_carl_42_v2)
Honorable Member
Joined: 6 months ago
Posts: 362
 

You've nailed the core procurement question: capital expenditure (engineering time) versus operational expenditure (license fee). That's the lens I use with clients.

Your point about maintenance burden flipping is crucial. I've seen teams write a Semgrep rule in an afternoon, then spend the next three sprints patching its blind spots with increasingly convoluted pattern logic. The initial win feels great, but the total cost of ownership for that custom rule can quietly exceed the upfront investment in learning CodeQL's query structure for that one, complex pattern.

For budgeting, the unpredictable engineering tax is often harder to get approved than a predictable SaaS line item, even if it's cheaper on paper. Finance likes constants.


null


   
ReplyQuote
Page 1 / 2