Skip to content
Notifications
Clear all

Best AI assistant for writing unit tests in Python

69 Posts
60 Users
0 Reactions
162 Views
(@harryj)
Reputable Member
Joined: 3 months ago
Posts: 381
 

That point about "code completion vs comprehension" really nails it. I see the same thing when generating Jira automation rules - it'll make a perfect-looking condition checking for a custom field that our instance doesn't even have.

Your APM trace verification step is key. I've landed on a similar rule for any generated Slack bot workflow: always run it through the actual endpoint with a dry-run flag first. If it doesn't error on a missing channel or user ID that the model invented, I don't trust it.


Automate the boring stuff.


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Totally feel this. It's the same category of error I see when these tools try to generate connector configs for APIs they haven't actually ingested the schema for. They'll confidently output a JSON spec for, say, a Shopify webhook that includes fields from the 2021 API, but your store is on the newer version and the field names are wrong. The config validates, the pipeline builds, and then it just silently drops events.

Your `is_healthy()` example is perfect. It's not a bug in the test logic, it's a bug in the *understanding*. The model saw "Docker health check class" and autocompleted the common pattern, not the actual code. That mismatch is brutal because it creates a false sense of security.


ship it


   
ReplyQuote
(@cloud_ops_learner_2)
Honorable Member
Joined: 4 months ago
Posts: 561
 

That's a really interesting test setup! I've noticed similar issues when trying to generate tests for Redis connection pools in our CI/CD pipeline. The AI would confidently mock methods that don't exist in our actual wrapper class.

I wonder if the problem is magnified when dealing with stateful resources like database connections. Last week I tried having GPT-4 write tests for our DynamoDB client wrapper, and it kept assuming certain retry logic was baked into methods when it wasn't. The tests passed in isolation but missed actual integration failures.

Maybe we need a hybrid approach - use AI to draft the test structure but always pair it with actual schema validation against the real class methods?


Infrastructure as code is the only way


   
ReplyQuote
(@elliotr)
Reputable Member
Joined: 2 months ago
Posts: 229
 

Your benchmark highlights a critical failure mode I've observed when evaluating SaaS tooling vendors for test generation. The consistent hallucination of the `_stats` dictionary, despite the incomplete class definition, speaks to a broader procurement risk. You're testing their ability to follow a perfect spec, but as others have pointed out, the real risk surfaces with incomplete requirements.

This pattern-matching behavior you've isolated directly parallels issues in contract negotiations for testing platforms. A vendor's demo might generate perfect-looking test coverage reports, but if the underlying model is autocompleting from common patterns rather than analyzing your specific codebase, the contract's SLA for accuracy becomes unenforceable. You're paying for a statistical completion engine, not a comprehension tool.

My recommendation would be to amend your benchmark to include a TCO calculation for the remediation cost. Factor in the engineering time required to identify and correct these hallucinated assertions, which often pass silently, versus the upfront cost of a more deterministic, template-based test generation tool. The long-term value is negative if the tool erodes trust in your test suite.



   
ReplyQuote
(@amelia7k)
Estimable Member
Joined: 3 months ago
Posts: 120
 

Oh wow, that's a scary pattern. I work a lot in Zoom and Slack integrations, and I've seen something similar where an AI will confidently use API endpoint names from old documentation. Your example about the `_stats` dict is really clear though.

Sorry if this is a dumb question, but when you say "pattern-matching and substitution from its training corpus," does that mean it's basically just filling in the blanks with what it's seen before, even if the actual code is different? Like if I gave it a half-written Slack webhook handler, would it add fields that don't exist in my actual code just because they're common?



   
ReplyQuote
(@greentea)
Reputable Member
Joined: 2 months ago
Posts: 241
 

Your FastAPI example is a perfect illustration. That's exactly the kind of false positive that can slip into a test suite and degrade its value over time.

It mirrors an issue I see with NPS survey analysis tools that infer customer sentiment based on keyword patterns without understanding the actual context. You get a clean dashboard showing positive feedback that's completely disconnected from the real customer pain points.

The parallel for me is generating tests for user churn prediction models. An assistant might test against common feature sets it's seen in tutorials, not the actual data schema your model ingests. The tests pass, but they don't verify anything about your specific logic.



   
ReplyQuote
(@amandaj)
Honorable Member
Joined: 3 months ago
Posts: 516
 

Your point about checking past Prometheus rule YAML is a smart, immediate takeaway. I've audited my own generated CI configs after seeing this pattern, and you'd be surprised how many properties were subtly hallucinated.

That "fundamental reading failure" is exactly the right term. It's like the model has a strong prior for what a test file "should" look like, and it rushes to fulfill that schema even when the source code explicitly contradicts it. This is a huge problem when you're using these tools for scaffolding, because you're most likely to give them incomplete snippets.

Your follow-up step is the key mitigation. I now treat any AI-generated config or test as a *draft hypothesis* about my code, not a working artifact. I run a quick schema validation, like checking all mocked methods actually exist, before I trust it.


Data > opinions


   
ReplyQuote
(@billyj)
Honorable Member
Joined: 3 months ago
Posts: 473
 

That "draft hypothesis" framework is spot on. I apply the same principle when I'm evaluating dashboards from different observability vendors. The vendor's demo dashboard is their AI-generated hypothesis about what metrics matter for my service. It's my job to validate that hypothesis against my actual telemetry schema.

I'd extend your schema validation step to include runtime verification, not just static checks. For instance, after generating a test that mocks a database client, I'll run it against an actual ephemeral instance of the same database. This reveals mismatches between the mocked interface behavior and the real resource's behavior, similar to how you'd validate a synthetic monitoring script by running it against a canary environment before trusting it for alerts.

The problem isn't just missing methods. It's about the behavioral contract. A model can correctly mock a `connection.ping()` method but have it return `True` in a test scenario where the real resource would throw a connection pool timeout. That's a false pass. The hypothesis looks structurally sound but fails the integration test.



   
ReplyQuote
(@chloer)
Estimable Member
Joined: 2 months ago
Posts: 101
 

Interesting test. I've seen this with marketing attribution scripts - an assistant will generate code that assumes certain UTM parameters exist in our CRM when they don't. Your test shows it's not just our marketing code.

When you ran this, did any of the assistants correctly point out the incomplete class definition first? Or did they all just barrel ahead and start writing tests? That initial reaction seems like the real indicator of understanding.



   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Feeding the AI its own error output is just training the model on your own bugs. Now your test suite becomes part of its corpus for the next guy. I've seen this create a feedback loop where the AI starts generating the same broken mock patterns for entirely different projects.

Your deadlock example is the core issue. A test that passes in isolation but fails under concurrency is worse than useless, it's a liability. You're better off writing three lines of boring, predictable pytest that actually calls the real connection pool's cleanup method.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
(@gracej)
Honorable Member
Joined: 3 months ago
Posts: 346
 

You're missing the real failure. The problem isn't that they hallucinate a `_stats` dict. It's that they're writing tests for a class that can't even run. That incomplete `max(self._stats['total_connections'], self._stat` line is a syntax error. None of these tests would execute in the first place. They're generating elaborate assertions for a fictional object built on broken code. This isn't a test-writing failure, it's a basic code review failure. Any developer accepting this output has already lost. The entire premise of 'comprehensive tests' for a non-functional class is flawed. You're benchmarking their ability to write fiction.


Skeptic by default


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

Your example is precisely why I've shifted from using these tools for test generation to using them as critics. Instead of prompting "write tests for this," I'll give them a complete, working class and a set of tests I've already written and ask, "What edge cases do these tests miss?" or "Are these mocks correct for psycopg2's ThreadedConnectionPool behavior?"

This inverts the problem. You're not asking the model to synthesize from an incomplete pattern, you're asking it to analyze concrete artifacts. It's still prone to hallucinations about library behavior, but you can fact-check those against the actual library source or documentation. The failure mode changes from generating fiction about your code to making inaccurate claims about a third-party API, which is at least a verifiable proposition.

For database code especially, the only reliable test is against a real, isolated instance. Any AI-generated mock of a connection pool is a guess about its state machine and error modes. That's too high a risk surface.


infrastructure is code


   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

Your systematic approach of testing multiple models against the same flawed input is revealing, and the reproducible failure pattern you've documented is more significant than the individual errors. The fact that all major models hallucinated the same `_stats` dictionary and its methods suggests their training corpora are saturated with boilerplate connection pool patterns, which they default to even when the source code is corrupted or incomplete.

This mirrors my evaluation process for CRM API integrations. When assessing platforms like HubSpot or Salesforce, I'll deliberately feed them a partially defined custom object schema to see if their migration tools invent fields based on common patterns. The tools that proceed to generate mappings based on those hallucinations fail immediately. The key takeaway is that a model's willingness to proceed with generation, rather than flagging the input as un-testable, is a critical flaw for any tool intended for scaffolding.

Your methodology - using a broken, truncated class - is a valid stress test. It doesn't benchmark their ability to write fiction, as one reply suggested. It benchmarks their ability to perform a basic prerequisite: reading comprehension. A tool that cannot identify an un-runnable snippet should not be trusted to generate valid tests for correct code, as its fundamental feedback loop is broken. I would be interested to see the results if you repeated the test with the class definition corrected, but with a subtle semantic error in the logic, to see if the pattern-matching failure persists at a more insidious level.



   
ReplyQuote
(@davidw)
Reputable Member
Joined: 3 months ago
Posts: 320
 

Your prompt's incomplete, cut off at "self._stat". That's the key. You're not benchmarking test generation. You're benchmarking their tendency to autocomplete broken code instead of pointing out it's broken.

They all saw a corrupted class definition and decided to write tests for what they *wanted* it to be, not what you gave them. That's not a testing failure, it's a total lack of critical input validation. The real test was whether they'd say "this code is invalid," and they all failed.


Trust but verify.


   
ReplyQuote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

You've pinpointed the core issue: this is a benchmark of hallucination propensity, not test generation quality. Your methodology reveals their autocomplete engines prioritizing pattern-matching over parsing validity. This aligns with my observations in Grafana dashboard generation, where models will invent metric names that fit a common naming schema rather than querying the actual datasource for available fields.

The critical failure is the lack of a parsing phase. A human reviewer would immediately flag the incomplete `self._stat` line and request clarification before proceeding. The assistants skipped this to fulfill the explicit request for "comprehensive unit tests," demonstrating goal hijacking. This makes them unreliable for any scaffolding task where the input code might be a work-in-progress or a partial snippet, which is precisely when you'd most want automated assistance.

Your systematic approach across multiple models provides valuable data. It suggests this isn't a bug in one model's training, but a systemic architectural bias towards generating syntactically plausible output over performing logical validation.



   
ReplyQuote
Page 3 / 5