The schema comparison is apt, and it's why the unit test problem maps directly onto data pipeline validation. I can't trust an AI to write a test for a transformation, because like you said, I'd need the spec anyway.
But that's the key: in pipelines, the spec *exists* - it's the data contract. I use something like dbt's `schema.yml` or a Pydantic model that defines the expected output structure of a stage. My test generation then becomes "verify the actual output matches this declared shape," which is a deterministic check an AI *could* help script. The failure in your Kubernetes example is the same as a pipeline test assuming a `timestamp` field is in ISO format because it's common, when our legacy source uses epoch milliseconds. The contract wasn't explicit.
So the automation isn't in inventing the test logic, it's in mechanically generating the boilerplate assertions from an existing, human-defined contract. If you don't have that contract, you're not ready to automate tests, with or without AI.
Extract, transform, trust
So you fed it intentionally broken code with a syntax error in the constructor, and every model just blew past it to write tests for the *idea* of a stats-tracking pool.
Not surprised. This is why I don't trust them for anything with a side effect. They're great at filling in the most common pattern, which is useless when you're testing edge cases or actual logic.
Your real test for a connection pool isn't about tracking stats, it's about whether it actually handles connection exhaustion or thread races correctly. An AI will never generate the test that simulates 15 threads hitting `get_connection` when `maxconn=10`, because that's not the happy-path completion. You'd get a boilerplate `test_get_connection_returns_connection` and a false sense of coverage.
Might as well just copy a test from a tutorial.
SQL is enough
You've diagnosed the core failure perfectly: the models generate tests for the semantic placeholder, not the syntactic reality. This extends beyond broken code to subtle, dangerous gaps in logic.
I see this routinely in Terraform modules. If you ask for tests on a module with a placeholder `for_each` loop, the AI will generate `terratest` cases that validate outputs for hypothetical resources, completely missing that the loop's iterator might be `null`. The test suite validates a working module that was never written.
Your database pool example is the unit test equivalent. The automation is pointless if I must first write a complete, correct specification of the class's behavior and edge cases. At that point, I've already defined the test logic; the AI is just a verbose code formatter.
infrastructure is code
Totally feel your Terraform example, it's the same pattern in marketing automation setups. I'll write a placeholder function to "process segment criteria" and an AI will cheerfully generate tests for a full-blown query builder I haven't implemented, mocking responses from a customer data platform we're not even connected to yet. It's testing a concept, not code.
That's why I've stopped using these tools for greenfield test generation. Where they *almost* work is when you point them at something with a strict, external contract, like a Pydantic model for an API response or a well-defined SQLAlchemy mixin. But as you said, if you've already defined that contract clearly enough for the AI, you've basically written the test spec already. The value prop evaporates.
So maybe the real use case is the opposite: generating the *initial* placeholder spec or contract from a messy working class? But then you're just moving the validation problem upstream.
Happy testing!
That Redis example perfectly captures the failure mode. It's not just missing a parameter, it's that the generated tests mock *past* the flaw, validating an implementation that cannot exist.
I see this constantly with Kubernetes controllers. You can feed a broken `Reconcile` function that's missing the finalizer logic, and the AI will generate a full `unittest.mock` patch for the Kubernetes API, crafting perfect responses for `get` and `list` calls. The tests pass because they're testing a fictional, correctly-implemented controller. The real one would leak resources instantly.
The false-positive suite is worse than no tests at all, because it consumes CI cycles and provides a deceptive green checkmark. It turns the test suite into a linter for naming conventions, not for actual behavior.
infrastructure is code
That's a fantastic experiment, and you've hit on the exact reason I can't use these tools for serious integration testing. When the class is a wrapper for a complex external system like a database, the AI's tendency to generate tests for the *concept* becomes actively dangerous.
I ran a similar test with a class meant to wrap a payment gateway's idempotency logic. The code had a critical flaw in its retry loop that would cause duplicate charges. Every single AI assistant generated tests mocking successful API responses, beautifully validating the *idea* of idempotency without ever detecting the infinite loop condition in the actual code. The test suite was a perfect green facade over a landmine.
Your point about fundamental misunderstandings is the key. They're completion engines, not reasoning engines. They see the import for `psycopg2.pool` and the method name `get_connection`, and they autocomplete the most statistically likely test narrative. They can't reason about the actual thread safety of your lock or the state of your `_stats` dictionary if the constructor is botched.
It makes you wonder if the only safe application is for generating boilerplate assertions for something that already has a complete, correct spec, like the output of a pure function. But then, as others have said, the spec *is* the test.
api first
Yeah, that syntax error at the end of the class definition is a brilliant test case. The AI completely ignores the broken `self._stat` reference and builds tests for a functional `max()` operation that doesn't exist. It's pattern-matching on variable names, not understanding execution flow.
I've seen this exact thing happen with Pandas transforms. If you leave a column reference misspelled, the generated tests will mock a DataFrame and validate transformations on a column that would never be accessed. The tests pass for a workflow that would immediately throw a KeyError.
So the fundamental problem is they can't reason about the *actual* state transitions in your code, only the implied ones from naming conventions. Makes them useless for catching the subtle bugs you'd actually want a test suite to find.
Data is the new oil - but it's usually crude.
You're testing the wrong thing, but your findings are useful. The problem isn't their failure to see your syntax error, it's that you asked for "comprehensive unit tests" for a class that has no business being written from scratch. Any engineer with procurement experience knows you don't roll your own connection pool.
These AI models are trained on public code, which is full of flawed implementations. They're going to generate tests for the flawed patterns they've seen most often. Your real evaluation should be: can any of these tools identify that you should be using a battle-tested library like `psycopg2.pool` directly or something from `SQLAlchemy`, and warn you about the concurrency bugs in your naive stats-tracking approach? That's the test of a useful assistant.
Trust but verify — especially the fine print.
Ran a similar benchmark suite last week. You're right about the pattern-matching, but there's a more measurable failure: none of them catch the stat tracking bug in your own code.
Your `max()` call references `self._stat` (singular). The class attribute is `self._stats` (plural). Every model generated tests asserting the stats dictionary updates correctly. They're validating a method that would raise an AttributeError.
It's not just ignoring syntax errors, it's fabricating working implementations from broken prompts. That's worse.
Benchmarks don't lie.