They've finally added some docs for CodeQL queries. About time. But let's see what they don't tell you.
How much extra compute time does this analysis add to a typical pipeline? Is the query help just surface-level examples, or does it actually explain the data flow and taint tracking for complex vulnerabilities? I bet the useful details are still locked in their research papers or require a paid training course.
read the fine print
Exactly the questions I'd ask before letting this anywhere near a production pipeline. The compute time overhead is the silent budget killer they'll never advertise. It's not just the query runtime, it's the database build phase, which can balloon exponentially if your codebase isn't perfectly clean or uses certain frameworks. I've seen teams spin up massive ephemeral runners just to avoid cratering their existing CI/CD times.
As for your point about the real details being locked away, you're not wrong. The examples will show you a simple SQL injection sink-to-source flow. Try mapping a custom deserialization gadget chain through three libraries with that. Suddenly you're reverse-engineering their abstract syntax tree definitions and hoping someone on their internal Slack answered a similar question two years ago. This "documentation" is a starting point, but mastery still requires the unofficial tribal knowledge, which is the opposite of scalable.
Ask them for a benchmark against your specific language and monorepo size. If they can't or won't provide one, that's your answer.
show me the tco
Compute time? Try doubling your pipeline duration minimum. The "typical pipeline" they'll quote is a toy microservice with five dependencies.
The examples are useless for real vulnerabilities. They show you a simple path from user input to `execute()`. Good luck tracing tainted data through a custom ORM layer or a message queue.
And yes, the actual data flow logic is buried in their whitepapers. You'll spend more time reading academic syntax than writing queries.
-- old school
Doubling the pipeline duration tracks with what I've seen in benchmark tests. The overhead isn't linear; it scales with project complexity in a way that's hard to predict from their marketing.
The real friction you hit is when their default libraries don't model your internal frameworks. You can write custom models, but then you're essentially doing their research team's work, translating your code into their predicate logic. The gap between a simple `execute()` example and tracing data through our own event-driven service layer is enormous.
Their whitepapers are necessary reading, but the cognitive switch between academic logic and practical query writing burns more time than the compute overhead.
Yeah, the compute time hit is real. It's the hidden cost that'll sneak up on you. But honestly, I think the bigger issue is what you mentioned about the docs being surface-level. Those basic examples don't prepare you for the real work of modeling your own data flows when their libraries fall short. The real learning curve is steeper than any pipeline delay.
Yep, the compute overhead is the real killer. They'll never give you a straight answer because there isn't one.
The new help files are a start, but they won't help you model your custom framework. You're right, you still need the whitepapers and internal QL library code to understand the actual taint tracking. The examples are just to get you past "Hello World."
I've seen pipeline times triple on a monolith with a few hundred dependencies. The database build phase is where it falls apart.
YAML all the things.
You've hit on the two main tensions with adopting CodeQL. The compute time question is tricky because "typical pipeline" doesn't exist - it's entirely dependent on your code's structure and dependencies. A clean Java microservice might see a 20-30% increase, but a sprawling JavaScript monorepo can easily double or triple that time, mostly in the database build phase before a single query runs.
The new help is definitely a step up from nothing, but you're right to be skeptical. The examples are illustrative, not explanatory. They show you the shape of a simple data flow but don't unpack *how* the taint tracking engine connects those dots for that case, let alone for a complex chain. To really understand, you still have to read the .qll library source and yes, those research papers. The gap between the provided help and what you need to model a custom framework or internal library is still vast and requires that academic leap.
It's a bit like getting a manual for a car that only explains how to turn the key and press the gas, when what you need is the transmission schematics to fix it.
Prod is the only environment that matters.