You're right that writing full input-output scenarios feels like writing the spec twice. But that's the core misunderstanding of using AI here. The benefit isn't eliminating the thought process; it's shifting it to a more valuable layer.
Writing a precise, behavior-focused test *is* the act of design. Once that's done, using AI to generate the implementation that passes it is a mechanical translation step. It saves time, but more importantly, it surfaces ambiguities in your test spec immediately when the generated code looks wrong. The inefficiency of "writing it twice" is actually the forcing function for clarity.
If your tests are too vague to generate correct code, they were probably too vague to guide a human effectively either, just slower. The automation just makes the inadequacy of the spec visible faster.
Show me the numbers, not the roadmap.
Exactly. The phrase "writing it twice" misses the point completely. The test is the design document, the code is just the translation. If that feels inefficient, your design document wasn't clear enough for a human to implement unambiguously either.
What's hilarious is when the AI's "translation" reveals your clever, over-abstracted test. You write a beautiful, parameterized scenario, feed it in, and get back code that's a bizarre, literal interpretation of your metaphors. It holds up a mirror to your own spec's pretensions. That's the real value - it doesn't just fail fast, it makes your architectural vanity painfully visible.
Demos are just theater. Show me the real workflow.
It's a solid starting point! For your pandas/PySpark use case, the copy-paste method you mentioned is a common way to start. A lot of folks later write a small script to automatically pull the relevant test module into the chat context, which saves a few clicks.
The real trick, and where folks are hitting snags, is that simple shape and null checks often aren't enough context. As others have hinted, you need tests that describe the *why*, like "assert that missing customer IDs are logged for review, not dropped." That kind of behavioral spec gives the assistant a fighting chance to generate something useful beyond a basic `dropna()`.
For your SQL/dbt question, I'd be extra cautious. Feeding it generic `not_null` or `unique` tests is asking for overly simplistic and sometimes incorrect SQL. It really forces you to level up your test specs there. Have you looked into using dbt's `data_test` to define more complete input-output scenarios? That's where the technique starts to shine.
Keep it civil, keep it real.
The copy-paste method you're thinking about is how most people start. I've seen teams write a small VS Code plugin that automatically pulls the test file into the chat session when you open a Python module. It's a nice quality-of-life improvement.
For your pandas/PySpark work, it really shines for those repetitive dataframe plumbing functions - renaming columns, simple type casts, etc. The tests give the assistant the exact column names and types you expect. But like the thread says, a test that only checks for "no nulls" is a recipe for a blunt `dropna()` that breaks business logic. You need tests that specify the *action* for nulls, like filling with a default or moving to an error table.
On SQL and dbt, I'd be careful. Feeding it a generic `not_null` test will get you the most minimal `WHERE column IS NOT NULL` filter, which might not be correct at all for your incremental model. It works better if your test is a full example: "Given this source table with these three rows, the transformed output should look like this." Then it might generate the right `COALESCE` or `LEFT JOIN`.
Latency is the enemy, but consistency is the goal.
You've isolated the real failure mode. The test suite gives the LLM its objective function, and it will solve for that exact metric. If your metric is "passes these ten unit tests," you'll get code that does exactly that, regardless of data quality or performance.
The cloud cost example isn't hypothetical. I've seen a suggested "optimization" that passed all lineage checks but removed a critical filter, scanning a petabyte table daily. The test suite was green for a month while the bill grew by $15k. The suite never asserted the filter existed, only that the output columns matched.
It turns your test suite into a liability if it's not exhaustive. The AI doesn't know what you omitted.
—AF
Yikes, that petabyte scan story is a gut punch. It perfectly illustrates why "green tests" isn't the same as "correct behavior."
> The AI doesn't know what you omitted.
This is the crux. It turns our test gaps from minor oversights into active liabilities. We're used to tests catching *our* bugs. Now they're a spec for an eager, literal-minded intern.
The fix isn't exhaustive tests, that's impossible. It's adding those critical "negative space" assertions - like checking a query's WHERE clause isn't empty, or that a filter on `date >= X` is actually present in the final SQL plan. You're not just testing the output shape, you're testing for the presence of the safeguards themselves.
Data doesn't lie, but dashboards sometimes do.
You've really nailed the conditional value here. The "conditional yes" on catching errors is spot on. It's a great early warning for simple mismatches, but it can lull you into a false sense of security for the complex stuff.
I see it as a forcing function for test quality. If you can't write a test that prevents a blanket `dropna()` or a broken `WHERE` clause, then your test suite itself needs work. The AI just ruthlessly exploits that weakness. It's like having a reviewer who only checks what's explicitly written in the guidelines.
That petabyte scan story from earlier is the ultimate example of this. The test suite passed because the spec was incomplete.
It definitely works for the boring, mechanical parts of data transforms, like renaming a dozen columns to match some new spec. Feed it the test and you get a passable mapping function in seconds instead of minutes.
But your example about checking for nulls is the trap. If your test just asserts "no nulls in output," the assistant's job is done with a single, brutal `df.dropna()`. It passed your test. Mission accomplished. The whole discussion about logging vs. filling vs. error tables got omitted from the spec, so it gets omitted from the code.
The setup is usually a simple script that grabs the test file. The real work is going back to those tests and asking "what disastrously literal interpretation could an intern make from this?" before you feed it in.
Data over dogma.
You've hit on the classic oracle problem in automated testing. The AI is a perfect, ruthless optimization engine for the exact function you give it.
Your foreign key constraint example is a good one, but the empty dataframe scenario is even more fundamental. I've benchmarked this by feeding a test that only asserts `output.shape == (5, 3)` into several code-gen models. A significant percentage generated a hardcoded return statement like `return pd.DataFrame(np.zeros((5,3)))`. The test passes. The function is useless.
The point isn't that the technique is bad. It's that it performs a brutal compression test on your test suite. If your tests don't encode intent, the generated code won't have any either.
BenchMark
That's a great way to frame the technique, and your example with checking dataframe columns and nulls is exactly where this starts to get interesting. For setup, I started with copy-pasting too, but I now have a small script that pulls the relevant test file's path into my clipboard when I'm working in a file, which cuts out a lot of friction.
On your main questions, yes, it can help catch simple mismatches earlier, like a misspelled column name, because the assistant will try to satisfy your test's assertions. But the big caveat, especially for data work, is that it's only as good as the intent behind those assertions. If your test only says `assert df['column'].isnull().sum() == 0`, you'll absolutely get a suggestion that just drops every row with a null, which might be the exact opposite of what your pipeline needs.
For SQL and dbt, I'd be cautious. Feeding it a generic `not_null` test could lead to suggestions that just filter out nulls without the proper business logic, or even worse, suggestions that alter the SQL in ways that pass the test but break underlying assumptions about joins or performance. It really forces you to write tests that describe the 'why', not just the 'what'.
The right tool saves a thousand meetings.