Having recently undertaken a comprehensive evaluation of our GitHub Advanced Security (GHAS) implementation, I specifically focused on the extensibility offered by custom CodeQL query packs. While the out-of-the-box security queries are robust, our team sought to incorporate checks for domain-specific patterns and internal library misuse. This led us to explore the ecosystem of third-party query packs.
The immediate benefit is clear: expanded coverage without developing every query in-house. However, a methodical analysis revealed several trust and quality considerations that I believe warrant thorough discussion.
**Primary Trust Concerns:**
* **Provenance and Maintenance:** The origin of a query pack is critical. Packs from well-known organizations with clear governance are preferable to anonymous repositories. One must assess:
* Update frequency and commitment to supporting newer CodeQL versions.
* Transparency in the query-writing process (e.g., are PRs reviewed?).
* The license and any associated liabilities.
* **Query Quality and Precision:** A high false-positive rate from a third-party pack can cripple developer adoption. Key quality indicators include:
* The presence of explanatory metadata and precise `@problem` and `@precision` tags.
* Whether queries are tested against real-world codebases (e.g., OSS benchmarks).
* The complexity of the underlying data flow or taint-tracking logic; overly simplistic queries can miss variants or introduce noise.
**A Practical Example:**
We trialed a popular third-party pack for "Potentially Dangerous Function" detection in our JavaScript codebase. Initial runs flagged hundreds of instances. Closer inspection revealed the pack lacked context sensitivity, flagging all usage of `eval()` even within secure, internal tooling where arguments were fully controlled. We had to clone the pack and refine the taint steps, which defeated the "off-the-shelf" benefit.
```yaml
# Example of a problematic custom query structure we encountered
- description: "Uses eval function"
severity: error
# Missing: No source or sink definitions for taint flow.
pattern: |
callExpression(callee=/eval/);
```
**Recommendations for Evaluation:**
Before integrating any third-party pack into a production CI/CD pipeline, I advocate for a structured pilot:
1. **Isolate and Audit:** Run the pack in a reporting-only mode against a snapshot of your code. Manually review a statistically significant sample of findings.
2. **Measure Signal-to-Noise:** Calculate initial precision (`True Positives / All Findings`) and recall (if you have a known vulnerability set).
3. **Inspect Query Logic:** Examine a subset of the `.ql` files for soundness. Look for proper use of data flow libraries, sanitizer steps, and clear provenance tracking.
4. **Assess Performance Impact:** Some complex packs can significantly increase analysis time. Benchmark against your baseline.
My current stance is that third-party packs are a powerful augmentation, but they must be treated as untrusted code until validated. The due diligence process is non-trivial and often requires senior CodeQL expertise. I am interested in hearing from others who have established formal vetting processes or can recommend packs that have demonstrably passed a high bar for quality and maintenance.
— Amanda
Data > opinions
You're hitting on the crucial point everyone glosses over: the false positive rate. Even a pack from a prestigious origin can be toxic if it's not tuned for real-world codebases. I've seen teams adopt a "high signal" pack from a respected firm only to have their devs start ignoring all security alerts within a month because 80% of the findings were speculative or contextually irrelevant.
Your methodical approach is good, but I'd add one more litmus test: run the pack against a snapshot of your own, known-clean historical code. Don't just review the queries abstractly. If it flags a dozen "issues" in code that's been in production for five years without incident, that tells you everything about the author's assumptions versus your actual risk profile.
show me the tco
You're absolutely right about the real-world test. I've formalized that step into what I call a calibration audit before any pack adoption. The process is simple: run the candidate pack against a curated, versioned snapshot of your own code that represents a "known-good" baseline, often the last externally audited release.
This does more than just measure false positives, though. It reveals the pack author's threat model. If it flags a dozen items in your stable code, it often means they're prioritizing theoretical vulnerabilities over exploitable ones, or they're writing for a different tech stack maturity. That mismatch in risk tolerance is a deal-breaker, regardless of the source's prestige.
One caveat to your method: it assumes your historical code is genuinely clean. In my experience, teams are often surprised when a high-quality pack finds a latent, legitimate issue they'd missed for years. The key is to triage those findings meticulously. If it's a true positive on old code, that's a strong argument for the pack's value, not against it.
—at
Absolutely agree on the **Provenance and Maintenance** point. I've been burned by a pack from a seemingly reputable researcher that just stopped getting updates after a CodeQL language pack change. Our CI started failing silently because the queries were incompatible, and we lost coverage for a sprint before we figured it out.
One thing I'd add to the quality indicators: check if the pack includes any test code alongside the queries. A good sign is when authors provide `.qlref` files or example code snippets that are supposed to trigger the query. It shows they actually validated the logic against something concrete, not just wrote a theoretical pattern.
The license bit is huge, too. Saw a pack with a weird non-commercial clause buried in the docs - legal had a field day with that one.
Totally agree that a calibration run can cut both ways. I love the term "calibration audit", by the way.
Your point about a high-quality pack surfacing a *legitimate* latent issue is spot on. We had that happen with a pack focused on dependency confusion in internal package managers. It flagged a pattern in our artifact publishing scripts we'd written years ago and completely forgotten about. The finding was valid, and it became the strongest endorsement for adopting the entire pack.
The trick is not letting those "good" surprises bias you during the initial evaluation. I still think a high false positive rate against stable code is a red flag, even if it also finds one true gem. You have to weigh the signal-to-noise ratio for your team's capacity.
Prompt engineering is the new debugging
Provenance matters, but it's not a guarantee. I've seen packs from major vendors with queries that have huge blind spots because they're written for a generic audience.
Your last bullet is the real filter: precision. Don't just look at the query logic, check if they publish their precision and recall numbers from testing. If they don't, they likely haven't measured it.
A pack that doesn't document its false positive rate is a liability. You'll inherit their technical debt.
Five nines? Prove it.
You're absolutely right to focus on **precision** as a top filter. A high false positive rate just trains developers to ignore alerts, which defeats the whole purpose.
I'd add that checking for "test code" alongside the queries, like user1212 mentioned, is a great proxy for that precision. If an author hasn't built a simple test case to validate their own query, they likely don't know its real-world behavior. That's a major red flag for me.
One practical step we take is to fork any third-party pack we're evaluating. Before we even run the calibration audit, we scan the query files themselves for obvious quality issues: are there extensive `@precision` tags? Are the `@problem` messages helpful and actionable, or vague and scary? It's a quick way to filter out the low-effort packs.
Ship fast, measure faster.
I like your practical step of forking and scanning the query files. That's a solid, hands-on check.
One thing I'd add to looking for `@precision` tags and helpful `@problem` messages: also check the `@description` metadata. A high-quality query pack should explain *why* the pattern matters and what the actual risk is, not just label it as "bad." If the description is just "Avoid this," it often means the author hasn't thought through the real-world impact, which correlates strongly with noisy, low-precision alerts.
Your point about test code being a proxy for precision is a great shortcut. No test files often means the author hasn't done the hard work of validation.
Keep it constructive.