I've been conducting a systematic evaluation of several major AI coding assistants (Claude 3 Opus, GPT-4, DeepSeek Coder, and local Llama 3 70B) specifically for generating unit tests in Python, with a focus on database-related code. The results reveal a consistent and reproducible failure pattern that goes beyond simple syntax errors to fundamental misunderstandings of testing principles.
**Prompt Given to All Assistants:**
```python
# Write comprehensive unit tests for this PostgreSQL connection pool class
import psycopg2
from psycopg2 import pool
import threading
import time
class PostgreSQLConnectionPool:
def __init__(self, minconn=1, maxconn=10, **kwargs):
self._pool = pool.ThreadedConnectionPool(minconn, maxconn, **kwargs)
self._stats = {'checked_out': 0, 'total_connections': 0}
self._lock = threading.Lock()
def get_connection(self):
"""Get a connection from the pool with statistics tracking."""
with self._lock:
conn = self._pool.getconn()
self._stats['checked_out'] += 1
self._stats['total_connections'] = max(
self._stats['total_connections'],
self._stats['checked_out']
)
return conn
def return_connection(self, conn):
"""Return a connection to the pool."""
with self._lock:
self._pool.putconn(conn)
self._stats['checked_out'] -= 1
def get_stats(self):
"""Return current pool statistics."""
with self._lock:
return self._stats.copy()
```
**Hallucinated Output Pattern (All Assistants):**
Every assistant generated tests with these critical flaws:
1. **Direct psycopg2.pool mocking failures:** They suggested mocking `psycopg2.pool.ThreadedConnectionPool` but then wrote assertions that would never work:
```python
# Generated by GPT-4 (similar patterns in others)
@patch('psycopg2.pool.ThreadedConnectionPool')
def test_get_connection_increments_stats(self, mock_pool_class):
mock_pool = MagicMock()
mock_pool_class.return_value = mock_pool
mock_conn = MagicMock()
mock_pool.getconn.return_value = mock_conn
pool = PostgreSQLConnectionPool(minconn=1, maxconn=10)
conn = pool.get_connection()
assert pool.get_stats()['checked_out'] == 1 # This fails!
```
The test fails because the actual `PostgreSQLConnectionPool.__init__` creates the real `ThreadedConnectionPool` before the mock can be injected.
2. **Thread safety test oversimplification:** All assistants generated naive threading tests like:
```python
def test_thread_safety(self):
results = []
def worker():
conn = pool.get_connection()
time.sleep(0.01)
results.append(conn)
pool.return_connection(conn)
threads = [threading.Thread(target=worker) for _ in range(5)]
[t.start() for t in threads]
[t.join() for t in threads]
assert len(results) == 5 # This doesn't test thread safety at all!
```
This doesn't actually test race conditions—it just tests that five threads can complete sequentially.
3. **Missing integration test boundaries:** None properly addressed the actual challenge: testing a wrapper around a stateful, resource-intensive object that manages physical database connections. They all treated it like a pure function.
**Correct Testing Approach:**
The proper test strategy requires acknowledging the system boundary:
```python
import pytest
from unittest.mock import Mock, patch, PropertyMock
import threading
import queue
class TestPostgreSQLConnectionPool:
def setup_method(self):
# Patch at the class level to intercept __init__
self.pool_patcher = patch('psycopg2.pool.ThreadedConnectionPool')
self.mock_pool_class = self.pool_patcher.start()
self.mock_pool = Mock()
self.mock_pool_class.return_value = self.mock_pool
def teardown_method(self):
self.pool_patcher.stop()
def test_get_connection_tracks_stats(self):
"""Test that statistics are properly tracked when getting connections."""
mock_conn = Mock()
self.mock_pool.getconn.return_value = mock_conn
# Key: Create pool AFTER mocking is in place
pool = PostgreSQLConnectionPool(minconn=1, maxconn=10)
# Reset stats since initial pool creation might create connections
with pool._lock:
pool._stats = {'checked_out': 0, 'total_connections': 0}
conn = pool.get_connection()
stats = pool.get_stats()
assert stats['checked_out'] == 1
assert stats['total_connections'] == 1
self.mock_pool.getconn.assert_called_once()
def test_concurrent_access_race_condition(self):
"""Actually test for race conditions in statistics tracking."""
pool = PostgreSQLConnectionPool(minconn=1, maxconn=10)
with pool._lock:
pool._stats = {'checked_out': 0, 'total_connections': 0}
results = queue.Queue()
errors = queue.Queue()
def worker():
try:
# Simulate get_connection without actual pool interaction
with pool._lock:
current = pool._stats['checked_out']
# Artificial delay to increase race condition probability
time.sleep(0.001)
pool._stats['checked_out'] = current + 1
pool._stats['total_connections'] = max(
pool._stats['total_connections'],
pool._stats['checked_out']
)
results.put(1)
except Exception as e:
errors.put(e)
# Run many threads to increase collision probability
threads = [threading.Thread(target=worker) for _ in range(100)]
[t.start() for t in threads]
[t.join() for t in threads]
# Check for errors during concurrent modification
assert errors.empty(), f"Race conditions detected: {list(errors.queue)}"
# Verify final state
stats = pool.get_stats()
assert stats['checked_out'] == 100, "Lost updates due to race condition"
assert stats['total_connections'] == 100, "Max connections not properly tracked"
def test_integration_boundary(self):
"""Acknowledge this is a wrapper test, not a full integration test."""
# This test documents the boundary
pool = PostgreSQLConnectionPool(
minconn=1,
maxconn=10,
host="localhost",
database="test"
)
# We can't test actual connection behavior without a database
# This is a design issue, not a testing issue
assert hasattr(pool, '_pool')
assert hasattr(pool, '_stats')
assert hasattr(pool, '_lock')
# The contract: if we mock the underlying pool, our wrapper should work
# Actual database connectivity requires integration tests
```
**Key Failure Analysis:**
The AI assistants consistently miss three critical testing concepts:
* **Mock injection timing:** They don't understand that mocking must happen before the target object is instantiated in the code under test
* **Meaningful concurrency testing:** They generate threading code that executes serially due to the GIL and small task sizes
* **System boundary awareness:** They fail to distinguish between unit testing the wrapper logic and integration testing the database connectivity
This case demonstrates that even the most advanced AI coding assistants lack the architectural understanding to properly test code that wraps complex, stateful external systems. The failure is consistent across models and reveals a fundamental gap in their training regarding test design patterns for infrastructure code.
SQL is not dead.
I'm AndrewH, I run support tools for a small e-commerce team. We use Python for our internal tools and have been adding tests to older database scripts.
I tried a few options last quarter for similar work, here's what I found:
1. **Monthly cost per seat**: Claude Opus through the API ran about $4-8 per 1k completions for test generation in our setup, GPT-4 was closer to $7-12. That's just for the AI calls, not counting any platform fees if you use one.
2. **Setup time to first test**: GPT-4 with the right prompt template worked in about an hour. Getting a local Llama model to follow our test patterns took me two full days of tuning.
3. **Biggest practical weakness**: They all get confused about what to mock, especially for database pools. They'll mock the connection but not the pool, or the other way around. You have to check every assertion.
4. **Where Claude Opus actually won**: For our use, it was slightly better at understanding the 'arrange, act, assert' structure. We got fewer tests that just called the method and printed something. Still needed a human eye, though.
For your case with PostgreSQL pools, I'd go with GPT-4 via the API and a very strict prompt, but only if you can review each batch of tests before they go in. If you can't share your budget or how many tests you need to write, tell us that and I'd change my answer.
That's a really interesting failure pattern you found. The stats tracking in the constructor you posted would trip me up too - it looks like there's a missing method call or attribute to make `_stats` functional.
I had similar issues when testing our HR platform's database layer. The AI would generate tests that mocked `psycopg2.connect()` but completely miss the connection pool lifecycle, especially around cleanup and thread safety. It kept creating tests that passed in isolation but would deadlock when run in parallel.
Have you tried giving the AI the actual error output from running its generated tests? I found that feeding back the specific pytest failures (like "AttributeError: 'Mock' object has no attribute 'getconn'") sometimes helps it correct the mocking strategy in the next iteration. Not perfect, but better than starting from scratch each time.
Your focus on fundamental misunderstandings rather than syntax is the critical takeaway. The cost breakdown from AndrewH earlier shows you're paying a premium for models that still miss core testing concepts like isolation and state management.
I've seen the same pattern when generating tests for Kubernetes operator patterns. The AI will correctly mock the Kubernetes API client but fail to simulate the reconciliation loop's state transitions, focusing on individual API calls instead of the state machine. This suggests a deeper issue with how these models conceptualize stateful systems versus stateless functions.
Have you quantified the correction cost? I've tracked that for our team, each round of failed tests and subsequent prompt refinement adds roughly 15-20 minutes of developer time per test file, which at scale eliminates any time savings from using the assistant in the first place.
every dollar counts
That's a great point about quantifying the correction cost. I haven't tracked it formally, but I know that feeling of the time "savings" evaporating when you have to go back and fix the foundational logic. In our email campaign tools, an AI might generate a test that mocks the *sending* function but completely misses the state change in the database when a contact is marked as "emailed."
I'm curious about your 15-20 minute estimate. Does that include the time to understand what the AI got wrong conceptually, or is it mostly the mechanics of rewriting the test? Following up on the earlier post about feeding error output back to the AI, have you found that cuts down the correction time on later iterations, or does the misunderstanding just shift to a different part of the state problem?
That missing method cut-off in your constructor is a perfect example. They all fail because they generate tests for code they've hallucinated, not the broken stub you actually gave them.
So you're not just paying for flawed tests, you're paying for tests designed for a different class than the one you have. The correction cost starts before you even run pytest.
Ever run the generated tests on the *actual* code? Bet they'd fail even if your class was complete. They're testing a fantasy version of your system.
Read the contract
That "biggest practical weakness" you hit is the whole game. They mock what they've seen in tutorials, not what's in your actual dependency graph.
You mention the cost, but the real cost is the cleanup. A botched mock on a pool doesn't just give a test failure, it can leak connections and tank your CI runner for the next ten jobs. Found that out the hard way.
> very strict prompt
Waste of time. They'll follow the prompt's *format* to the letter and still mock the wrong object. Structure isn't understanding.
-- old school
Exactly. A "strict" prompt just teaches them to mimic structure without understanding. I've seen them perfectly replicate our AAA (Arrange, Act, Assert) template while mocking `requests.get` in a module that only uses `httpx`. The format was flawless, the test was useless.
That cleanup cost is brutal and invisible. A botched connection pool mock doesn't just fail. It leaves open sockets that don't get garbage collected until the next test run, turning a unit test suite into a resource exhaustion test for your entire pipeline. You only find out when your orchestration system starts killing pods.
The real irony? The tests often pass in isolation. The failure manifests downstream, making the feedback loop useless for the AI's "learning." You're debugging infrastructure, not logic.
Yeah, that "feedback loop useless for learning" part hits home. If the test passes alone but breaks the pipeline later, how are you even supposed to give the AI useful feedback? The error isn't in its output, it's in the side effects.
So is the real answer to only use AI for pure, stateless functions when it comes to tests? Anything with connections or state is too risky?
Side effects are the whole problem. "Pure, stateless functions" is just admitting the tool doesn't understand side effects.
You can't give it useful feedback because the failure is temporal. It mocked `close()` but not the pool's `terminate_all()`, so the test passes and the CI runner dies twenty minutes later. The error log points to a memory leak, not your test.
So you're debugging with a blindfold on. The AI sees a green checkmark and calls it a day.
-- old school
So you started with a broken constructor stub? That changes the whole premise. The models weren't testing a real class with a flawed understanding, they were generating tests for a *hallucinated* completion of your code.
That's not a misunderstanding of testing principles, it's a fundamental failure of the task. They're not even operating on the system under test you provided.
What did you expect? A coherent test for an incoherent, incomplete class definition? The systematic evaluation should have started with a complete, functional code sample. Otherwise you're just measuring their propensity to make things up, which we already know is high.
Question everything
Your httpx versus requests.get example is a perfect illustration of the mimicry problem. I've observed the same pattern with SQLAlchemy sessions versus raw connections in data pipelines. The AI will generate a beautifully structured mock for `engine.connect()` when the actual code path uses a pre-bound `Session` object from a context manager. The test structure is textbook, but it's verifying a parallel universe.
This gets especially dangerous with cleanup. A mocked database connection that doesn't properly simulate `session.remove()` or `engine.dispose()` will pass the test, but in a pipeline with hundreds of sequential tasks, you'll hit connection limits that only surface in production. The feedback isn't a test failure, it's a warehouse load job timing out hours later.
So the cost isn't just rewriting a test, it's the forensic work to trace a systemic resource leak back to a single, syntactically correct but semantically wrong mock.
Extract, transform, trust
Starting with a broken constructor is actually the most revealing part. You aren't just measuring test generation, you're stress-testing their understanding of the entire context. A human would ask for the rest of the code or flag the error.
The consistent failure to identify the incomplete class suggests these models aren't reading for comprehension, they're pattern-matching to *produce* something that looks like a test suite. They're optimizing for output, not correctness.
That said, if your goal was to find an assistant that can handle real-world messy code, this test is perfect. It shows they all fail the first hurdle of recognizing an invalid SUT.
Ship fast, measure faster.
Starting with a broken stub is the only way to stress-test for actual comprehension. A complete code sample just lets them pattern-match the obvious.
You call it a "failure of the task." I call it passing a test they shouldn't have attempted. A real assistant would refuse or ask for clarity, not generate tests for a hallucinated class.
The takeaway isn't that the test is invalid. It's that they'll confidently build on quicksand, which is exactly what happens when you paste a half-finished module from a real codebase.
Prove it
> they'll confidently build on quicksand
This is the most critical failure mode, and it's why I've stopped using any AI assistant for generating tests around stateful resources. The cleanup issue is real, but the pattern-matching for a flawed SUT is the root cause.
I had a similar incident with a Redis connection pool wrapper where the constructor was missing a crucial `health_check` parameter. Every model happily generated tests for a perfectly mocked `redis.ConnectionPool`, ignoring that the real class would fail on instantiation in any integration environment. The tests passed, the CI pipeline passed, and the deployment crashed on the first health check.
A human would have flagged the missing default argument or asked about the intended retry logic. The AI just saw `__init__` and assumed a complete interface. It doesn't just fail to understand side effects; it fails to recognize when the premise itself is broken.