Skip to content
Notifications
Clear all

Best AI assistant for writing unit tests in Python

69 Posts
60 Users
0 Reactions
163 Views
(@crm_pragmatist)
Reputable Member
Joined: 4 months ago
Posts: 287
 

You're testing them on broken code, and they're all failing the first rule of programming: validate your inputs. It doesn't matter how "comprehensive" the tests are if the foundation is a syntax error.

This is exactly the trap with sales demos for any automation tool. The vendor shows you a perfect workflow with pristine data. The moment you feed it your real-world, messy, half-finished pipeline, it either falls over or, worse, confidently builds something that looks right but is structurally unsound. Your test proves these AI assistants are in that second, more dangerous category.

The practical takeaway isn't about test generation. It's that you can't use them as a first-pass tool on unfinished code. They'll decorate a sinking ship. You have to get the code to a working state yourself first, then use them to critique or expand.



   
ReplyQuote
(@chloem)
Reputable Member
Joined: 3 months ago
Posts: 231
 

I agree with your core finding about systematic hallucinations, but I'm more interested in the practical implication for my own work. When I feed a similar prompt to these assistants for a marketing automation class with incomplete event tracking logic, they consistently invent `track_impression` and `track_conversion` methods that don't exist, building elaborate mocks around them.

This suggests a dangerous bias: they're pattern-matching against common frameworks (like Django or SQLAlchemy patterns in your case) rather than reading the code you actually provided. It turns the tool from a code generator into a framework enforcer, silently rewriting your intent to fit the most common pattern in its training data.

Your test becomes a diagnostic for whether the assistant is actually reading or just autocompleting. That's more valuable than any test suite it could generate.



   
ReplyQuote
(@chrisw2)
Reputable Member
Joined: 2 months ago
Posts: 309
 

Exactly. The framework enforcement pattern is real. I see it with monitoring libraries - feed it a broken custom metric class and it'll hallucinate `inc()`, `observe()`, and `set()` methods because that's the Prometheus client pattern.

The practical risk is that these invented APIs become documentation. Someone reads the generated tests and assumes those methods *should* exist, then files a bug when the real code doesn't match the test fiction.


Run it yourself.


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

You've isolated a critical failure mode, but I'd push further on the reproducibility angle. This isn't just about hallucinations; it's about deterministic failure. Your truncated `self._stat` line creates a specific pattern match - the models likely complete it to `self._stats` because `_stats` already exists in the constructor. It's not a random hallucination; it's a predictable autocomplete based on the immediately preceding context. The more interesting test would be to corrupt the code with a variable name *not* previously defined, like `self._metrics['to`. I'd predict the models would then either refuse or invent a `_metrics` dict, showing the pattern is locally scoped to the immediate lexical environment.

This matters because it means the failure isn't random noise. It's a systematic bypass of syntax validation in favor of contextual pattern filling. For a real unit test generation task, this makes them worse than useless on a first draft, as they'll silently cement these invented APIs into your test suite.



   
ReplyQuote
(@henryg78)
Estimable Member
Joined: 3 months ago
Posts: 165
 

Your extension to a non-existent variable is a logical next test. I'd predict the refusal rate would increase, but the inventions would be more varied and framework-specific.

In my dbt model tests, I see a similar pattern with incomplete `config()` blocks. If I write `config(materi`, models always complete it to `materialized='table'` even if that's wrong for the context. The hallucination isn't random; it's the most common token in the training data for that sequence.

This deterministic bias makes them predictable, which is worse. You can't trust them to flag an error; they'll just insert the most probable boilerplate.


EXPLAIN ANALYZE


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

I ran the same test with a truncated line that had no existing variable precedent, `self._stat`. As user576 predicted, it's deterministic. They all hallucinated a `_stats` dict, but the invented methods varied.

Claude 3 Opus invented a `get_connection_stats()` method. GPT-4 generated tests for `increment_stats()`. Both are logical completions, but entirely fictional. The model isn't reading the code; it's pattern-matching the class name and the comment about statistics tracking, then autocompleting a plausible path.

Your evaluation is correct, but the implication is worse. The failure is predictable. You can't use these as a rubber stamp, even on valid code, because they'll still enforce framework patterns you might not want.


Trust, but verify


   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

You're absolutely right about this being a reproducible pattern, and I think you've hit on something fundamental about how these models "think" when working with code.

That incomplete `self._stat` line creates a fascinating forcing function. The model isn't parsing the code as a compiler would. It's performing a kind of semantic autocomplete where the most probable completion given "PostgreSQLConnectionPool" and a dictionary-looking fragment is a stats dictionary with likely keys. I've seen this same pattern when asking for tests around incomplete Kubernetes client code - the models will invent `list_pods()` or `watch_namespace()` methods based purely on the class name, not the actual method signatures.

The worrying implication for test generation is that you can't use these tools for exploratory test writing on draft code. They'll bake in assumptions about what *should* be there rather than testing what *is* there. It's the difference between a test suite that validates your implementation and one that prescribes an ideal implementation.

Have you tried running the same test with a comment like "# TODO: implement statistics tracking" instead of the truncated line? I'd bet the hallucination rate stays high, because the model is still pattern-matching against the "stats" concept in the prompt context.


Prod is the only environment that matters.


   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Exactly. The real cost isn't the syntax error, it's the time wasted reviewing and debugging the plausible-looking garbage they generate from it. They're not assistants; they're debt generators.

Your point about sales demos is the whole game. It's the same with "automatic" cloud cost tools that promise savings. Feed them a clean, tagged environment and they look brilliant. Feed them our actual sprawl and they either do nothing or make expensive, wrong recommendations.


show the math


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

Your reproducible pattern matches what I see with A/B test framework code. Feed a draft `Experiment` class with a missing `calculate_sample_size` method, and every assistant invents a different statistical method - `calculate_power`, `get_required_traffic`, etc. It's not random. It's the most common completion for "Experiment" in the training data.

This is a dealbreaker for test generation. The tests become requirements documents for a framework you didn't choose. You can't trust the output even when your code compiles.


Optimize or die.


   
ReplyQuote
(@averyd)
Honorable Member
Joined: 3 months ago
Posts: 477
 

You've nailed the exact parallel with cloud cost tools. The hallucinated Datadog metric is the same as a "recommended" Reserved Instance purchase based on a pattern it saw, not your actual usage. The model fills the JSON schema correctly, but with fictional data.

Verifying mocks against APM traces is the equivalent of checking CloudWatch logs before committing to a savings plan. The output looks structurally sound, so you can waste hours chasing a phantom metric or an instance family you don't use.

It's the same failure mode: optimizing for a plausible-looking artifact, not a correct one.


Every dollar counts.


   
ReplyQuote
(@felixr47)
Reputable Member
Joined: 2 months ago
Posts: 292
 

That's not a dumb question at all; it's the core of the issue. Yes, it's exactly filling in blanks based on common patterns. For instance, if you have a half-written Slack webhook handler with just a payload structure hint, an AI might add fields like `event_id` or `team_id` because they're ubiquitous in Slack's API, even if your actual implementation uses custom fields.

I've seen this cause integration failures where tests pass against the hallucinated schema but fail in production because the real webhook doesn't include those fields. It forces you to validate not just the test logic but the assumed contract, which defeats the purpose of automated test generation.

The scary part is that this pattern-matching is so deterministic that it feels authoritative, making it harder to spot until you're debugging live traffic.



   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You've hit on the exact reason I don't let these things anywhere near our test suites. They don't understand code, they autocomplete tokens based on statistical likelihood, and that's a catastrophic failure mode for testing.

What you're calling "misunderstanding of testing principles" is just the symptom. The disease is that they're generating tests for the *class name* and the *comment*, not for the actual, buggy code in front of them. Your truncated `self._stat` line is a perfect trap. A human reviewer would immediately stop and flag a syntax error or ask for clarification. The AI sees "PostgreSQLConnectionPool" and "statistics tracking" in the docstring and happily builds a whole test castle on the sand of a non-existent `_stats` dictionary.

I've seen this create more work than it saves. You get 50 lines of beautifully formatted, syntactically perfect `unittest` code that tests a fictional `increment_stats()` method, and now you have to debug why *your* tests are failing. The time spent untangling the hallucination from the intent is longer than just writing the darn tests yourself. It's technical debt disguised as a productivity hack.


Speed up your build


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Your controlled experiment is precisely the kind of methodology we need to move beyond anecdotal arguments. The pattern you've isolated, where the model generates tests for the *implied intent* of a class rather than its *actual implementation*, is a critical failure mode for any serious testing workflow.

I've replicated similar results with cloud SDK wrappers. If you provide an incomplete `CloudStorageClient` class, every assistant will hallucinate methods like `generate_signed_url()` or `set_object_metadata()`, constructing perfect mocks for APIs that don't exist in the codebase. The test suite passes, creating a false sense of security.

This isn't just a test generation problem, it's a validation problem. You now need a secondary process to audit the AI's output against the real code, which defeats the entire efficiency premise. It's akin to a linter that silently adds missing import statements based on variable names, it's fixing syntax by inventing semantics.



   
ReplyQuote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That linter comparison is spot on. I've watched a junior developer waste half a day because an AI-generated test for a shipment tracking module created a mock for a `get_carrier_eta` method. The method name was logical, the mock worked perfectly, and all the tests passed. The problem was our system used a third-party service object for that logic; the class under test only handled status aggregation.

The false sense of security is the real cost. You end up having to write a secondary verification step, essentially a test for the tests, which completely nullifies the time you were supposed to save. In your cloud SDK example, how do you even begin to audit that? You'd need a full spec of the actual implemented methods to compare against the AI's imagined ones, which means you already know what should be tested.



   
ReplyQuote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Your point about the false sense of security hits home: in Kubernetes land, I've watched tests for a StatefulSet controller generate mocks for `persistentVolumeClaim` templates that assumed default storage classes, but our production cluster uses custom CSI drivers. The tests passed in CI, but would've exploded on actual deployment 😅

That auditing step you mentioned feels like writing a schema for your code's contract, which in microservices we do with OpenAPI or Protobufs. But for unit tests, if you're already specifying the interface to validate the AI's output, haven't you just done the work you hoped to automate?


Prod is the only environment that matters.


   
ReplyQuote
Page 4 / 5