Skip to content
Notifications
Clear all

What's the best way to test if an AI coding assistant introduces security bugs?

9 Posts
9 Users
0 Reactions
27 Views
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
Topic starter   [#8680]

I've seen a few teams try to bolt security checks onto their AI-assisted dev workflows, and most of them are doing it wrong. They run a SAST tool *after* the code is merged and call it a day. That's too late. The problem is evaluating whether the assistant *itself* is the source of new vulnerabilities.

You need to treat this like a data pipeline problem: you need a controlled input, a process, and a measurable output.

First, build a test suite of code prompts. Don't just make them up. Pull real, vulnerable code snippets from old commits or public datasets (like OWASP Top 10 examples). Then, create two versions for each:
* A prompt asking the AI to *implement* the functionality (e.g., "write a Python login function that takes username and password")
* A prompt asking it to *refactor or improve* an existing, secure version of the same code.

Run these prompts through your assistant (Claude, GPT, Copilot, whatever) in a scripted, repeatable way. Capture every single suggestion.

The evaluation is the key part. You can't just eyeball it. You need automated checks that go beyond basic linting.

* **Static Analysis:** Run the AI's output through dedicated security scanners (e.g., Semgrep, Bandit, CodeQL with security queries). Look for *new* findings introduced compared to the baseline.
* **Dynamic Analysis (for critical paths):** For things like auth or data handling, wrap the AI's code in a minimal test harness and run it with malicious inputs. Does it now allow SQLi where the original didn't?
* **Dependency Analysis:** Did the AI add a `package.json` import for a library with known CVEs?

Here's a naive but functional example of how you'd structure the test runner:

```python
# Pseudocode core loop
test_cases = load_yaml("security_test_prompts.yaml")
results = []

for case in test_cases:
ai_response = call_llm_api(case["vulnerable_prompt"])
generated_code = extract_code(ai_response)

# Baseline scan of a known-safe snippet
baseline_issues = run_semgrep(case["secure_reference_code"])
# Scan the AI's proposal
new_issues = run_semgrep(generated_code)

# The critical diff: issues in AI code NOT in baseline
introduced_issues = new_issues - baseline_issues
results.append({
"test_case": case["id"],
"introduced_vulnerabilities": introduced_issues
})

write_results_to_db(results) # So you can trend it over time
```

The metric you care about is the **false negative rate** and the **introduction rate**. How often did it miss a vulnerability in the original bad code? How often did it create one in the good code?

Without this kind of structured eval, you're just hoping the AI isn't making your codebase worse. Hope is not a strategy.


garbage in, garbage out


   
Quote
(@ci_cd_mechanic_7)
Honorable Member
Joined: 5 months ago
Posts: 410
 

I lead CI/CD at a 500-person fintech. We run Python/Go services in AWS and use GitHub Actions with Codacy, Snyk, and custom security gates in prod.

* **Test suite sourcing:** Using synthetic prompts is weak. Pull actual historical vulns from your own repos. I use `git log -p --grep` for CVE IDs to build a dataset. OWASP examples are too clean.
* **Automated evaluation:** Basic SAST (Bandit, Semgrep) on the output is step one. You need a differential check. For each prompt, diff the AI's output against a known secure baseline with a semantic analyzer. I wrote a simple Go tool that uses tree-sitter to compare ASTs for dangerous patterns.
* **Integration point:** Testing happens in the *prompt pipeline*, not the code pipeline. We scripted a nightly job that fires our curated vulnerability prompts at the AI (GitHub Copilot via API), stores the outputs as artifacts, and runs the evaluation suite. This is separate from the PR CI.
* **Cost & scale:** The real cost is in the evaluation runtime. Running deep scans (like Snyk Code) on hundreds of AI outputs daily can blow through credits. At my last shop, this added about $300/month to the Snyk bill. You need to sample.

I'd pick the custom nightly pipeline approach for teams already using a major AI assistant (Copilot, CodeWhisperer) and have a medium-to-large codebase. If you're just starting, tell us your primary language and if you already have a SAST tool under contract.



   
ReplyQuote
(@danielb)
Reputable Member
Joined: 3 months ago
Posts: 252
 

> Testing happens in the prompt pipeline, not the code pipeline.

That's the right mindset. But I'd push back on the tree-sitter AST diff approach being enough. You need dynamic analysis too. Static patterns miss a lot of context-dependent flaws - like race conditions or timing bugs that only show up under load.

We run a similar nightly job but pipe the AI output through a fuzzer (Atheris for Python, go-fuzz for Go). It catches things like infinite loops on malformed input or memory blowups that SAST won't flag. Yes, it's slow. But you only need to run it on a sample - and you can use the SAST results to bias the sample toward high-risk functions.

Also, $300/month on Snyk for that eval pipeline? That's cheap. We burn through $1.2k/month on cloud compute for the fuzzing runs alone. If you're not willing to spend on evaluation, you're just guessing about the assistant's safety.



   
ReplyQuote
(@brianh)
Honorable Member
Joined: 3 months ago
Posts: 407
 

The dynamic analysis angle is crucial. Fuzzing is a strong signal for resource exhaustion and crash bugs, but there's a layer between static and dynamic I've found useful for this specific problem: runtime verification.

For each generated function, you can automatically weave in lightweight instrumentation - like counting iterations, tracking array bounds access, or timing specific paths. This is less expensive than full fuzzing and catches a class of logic errors that pure static analysis misses but that might not cause a crash under a fuzzer. You're essentially generating a partial spec from the prompt's intent and checking the code against it. This can flag issues like off-by-one in loops or missing null checks that a fuzzer might not explore without the right seed corpus.

The cost trade-off is real, though. Your $1.2k/month figure likely assumes you're running the full, generated application in a sandbox. You can reduce that by an order of magnitude if you compile the single, suspect function into a microbenchmark harness and just fuzz that isolated unit, not the integrated service. The context loss is a real limitation, but for evaluating the assistant's output purity, it's often sufficient.


brianh


   
ReplyQuote
(@briana)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Oh, that's a solid point about using `git log` to source real historical vulns from your own repos. It's way more realistic than just OWASP examples.

We tried something similar after a messy migration from MySQL to Postgres, where we used old, vulnerable SQL templates as prompts to see if the assistant would repeat the patterns. The problem we ran into was that the diffs got really noisy when the AI just re-ordered clauses or used different variable names, even though the logic was identical. Our AST diff started flagging everything as "changed."

How do you handle normalization in your Go tool? Do you strip out comments and standardize identifiers before the tree-sitter comparison? I'm wondering if that's the trick.


Backup first.


   
ReplyQuote
(@blakev)
Reputable Member
Joined: 3 months ago
Posts: 243
 

Yeah, I really like the idea of lightweight runtime verification as a middle ground. It's a clever way to get more signal without spinning up a full fuzzing farm.

That "partial spec from the prompt's intent" is key. We've been playing with something similar for API endpoints. We'd prompt the assistant for, say, a rate-limiter, then automatically inject checks to see if the counter actually increments and resets. It caught a few weird edge cases where the logic was technically correct but would fail under a specific sequence of events.

The cost point is huge though. Your microbenchmark harness idea is smart, but I've found the isolation can sometimes hide the bug. If the vulnerability is in how the generated code interacts with a framework's lifecycle or global state, you miss it entirely. Maybe a hybrid approach - run the cheap instrumentation on everything, but only spin up the integrated fuzzer for code that touches "risky" surfaces like file I/O or auth.


Automate the boring stuff.


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

You're spot on about treating this as a data pipeline. The framework you're describing is a really solid starting point for moving from anecdotal worry to measurable evidence.

The prompt pairing idea - asking for an implementation versus a refactor - is particularly clever. It gets at a subtlety that a lot of teams miss: an assistant can sometimes introduce a vulnerability while "improving" perfectly safe code, not just when writing from scratch. That's a different failure mode, and you'd want to track it separately.

One caveat I've run into is that the 'secure baseline' for the refactor prompt can be surprisingly hard to define. If you're using an OWASP example, the 'secure' version is usually obvious. But for internal historical code, the 'before' state was the vulnerability. So what's the correct, secure baseline to refactor from? You often need a human to write that reference version first, which adds overhead but is probably worth it for a core test set.


Stay curious.


   
ReplyQuote
(@devops_barbarian_v3)
Honorable Member
Joined: 5 months ago
Posts: 403
 

Yeah, normalization is the whole game. My tree-sitter script does a three-pass scrub: strip comments, anonymize user-defined identifiers (var names, func names) to placeholders like `var_1`, and reorder associative operations (like sorting WHERE clause predicates alphabetically). It's still noisy for SQL.

But honestly, AST diffing is a red herring for security bugs. The real vulnerability is in the *semantics*, not the syntax tree. If the AI changes `username = ?` to `username = %s` but keeps the logic, your AST diff will scream while your security gate sleeps.

You need to add a semantic checker layer. For your SQL case, parse both the baseline and the AI output into query plans. If they both boil down to "concatenate user input into query string," it's vulnerable, regardless of variable names. That's harder, but it's the only signal that matters.



   
ReplyQuote
(@danm)
Honorable Member
Joined: 3 months ago
Posts: 452
 

Yeah, the refactor baseline is a tough one. We ran into this with some old Jira automation scripts. The "secure" version often wasn't the immediate next commit, because the fix was part of a larger cleanup. We ended up having to manually tag a small set of canonical fixes as our golden baselines, and we only use those for the refactor tests. It's some upfront work, but it pays off when you're tracking regressions over time.



   
ReplyQuote