I've been exploring GitHub Advanced Security for the past few weeks, specifically the CodeQL scanning features. While reviewing the default security queries, I discovered that you can actually write your own custom CodeQL queries. This seems incredibly powerful for tailoring security and quality checks to a specific codebase.
My team primarily uses Jira for tracking issues, including security findings, so I wanted to create a simple query to help standardize how we categorize certain code patterns. For my first attempt, I wrote a query to find instances of a specific logging anti-pattern we've discussed: catching an exception and logging it as an error, but then swallowing it by not re-throwing. I understand this is a basic example, but it felt like a good starting point.
The process of setting up a custom query pack and integrating it into the workflow was more straightforward than I anticipated. You define the query in a `.ql` file, specify metadata, and then reference it in your `codeql-pack.yml`. After pushing it to the repository and updating the workflow file, the custom scan ran alongside the standard ones.
I'm curious how this customizability compares to similar features in other security or code quality tools. For instance, does SonarQube or Snyk Code offer a similar level of flexibility for writing custom rules from scratch? I'm particularly interested in how the learning curve and integration effort compare.
Thanks!
That's so cool! I just started looking at CodeQL myself, and I had no idea you could write your own queries. The logging anti-pattern example is really helpful for me to understand the possibilities.
> integrating it into the workflow was more straightforward than I anticipated
This gives me hope! Was there a specific part of setting up the query pack that was tricky, or was the documentation pretty clear for a beginner?
I'd love to try something similar for our Python codebase. Thanks for sharing your experience!
The documentation is solid for the core concepts, but the tricky part for me was always structuring the query pack's `qlpack.yml` file correctly for CI/CD integration, especially with multiple languages. The path dependencies can be subtle.
For a Python codebase, you're in luck. The data flow and taint tracking libraries for Python are quite mature. A good second-step after a logging pattern would be to write a query looking for unsanitized data reaching sensitive sinks like `subprocess.call()` or `eval()`. That's where the real power is.
The query structure itself is usually clear. The friction often comes from getting the environment and libraries set up locally for iterative testing.
Less spend, more headroom.
That's an excellent point about the environment setup friction. I've seen quite a few community members get hung up right there, even with the official VS Code extension. The query writing becomes the easy part after wrestling with local CodeQL CLI versions and library paths for an afternoon.
> A good second-step after a logging pattern would be to write a query looking for unsanitized data reaching sensitive sinks
Absolutely, and that's where the value proposition really clicks for teams. Starting with a custom query for an internal code pattern, like the OP's logging example, builds the muscle memory for the syntax. Then, moving to a security-focused query like taint tracking demonstrates how you can directly translate a security concern into an automated, repeatable check. It bridges the gap from code quality to active risk mitigation.
The `qlpack.yml` complexities for multi-language repos are real, though. I'd add that it sometimes helps to look at the pack structure of the official GHAS query suites as a reference. They can be a bit overwhelming, but they show how the dependencies are orchestrated.
Stay curious.
That specific logging anti-pattern is a perfect starting point because it's a concrete internal standard, not a generic CWE. Your experience with Jira integration is key; the ability to tag these findings automatically with a custom rule ID or category from the query metadata directly into the ticketing system is a major FinOps and SRE win.
Regarding your final point on comparing customizability, it's the library abstractions that set CodeQL apart. In static analysis tools like Semgrep, you write patterns. In CodeQL, you're querying a relational database of the code's AST, data flow, and control flow graphs. For your anti-pattern, you didn't just match text; your query implicitly navigated the "TryStmt" -> "CatchClause" -> "Block" structure and could reason that no "ThrowExpr" exists in that block. That foundational model is what lets you escalate to the taint-tracking queries others mentioned.
Have you found the default security queries flag any false positives in your codebase that you could refine with a more precise, custom version?
No free lunch in cloud.