Skip to content
Notifications
Clear all

Guide: using your existing test suite as a live context source for suggestions

40 Posts
37 Users
0 Reactions
27 Views
(@data_pipeline_rookie_43)
Honorable Member
Joined: 5 months ago
Posts: 365
Topic starter   [#27909]

Hey everyone! I've been trying to level up my data pipeline game, especially around testing and making my workflows more robust. I read this cool idea somewhere about using your existing test suite as a "live context" for your coding assistant, like GitHub Copilot or Cursor, to get better suggestions.

Basically, instead of just asking it to write a function, you feed it your actual pytest files or unittest cases first. The theory is it learns the expected behavior and style from your tests, then suggests code that's more likely to pass them right away. Sounds almost too good to be true for cutting down on back-and-forth!

I'm super curious if anyone here has tried this with data engineering code. My typical work is in Python, doing ETL with pandas or PySpark, and I use pytest. For example, I have tests that check if my dataframe has the right columns after a transformation, or if nulls are handled correctly.

My main questions are:
* How do you actually set this up? Do you just copy-paste a test case into the prompt every time?
* Does it work well with SQL transformations, like ones you might test with a framework like dbt?
* Has it helped you catch errors earlier, or is it more of a neat trick?

I'm imagining this could be huge for orchestrating tasks in Airflow tooβ€”making sure your DAGs and operators meet their contracts. But I'm still a rookie, so I might be overthinking it 😅

Would love to hear from anyone who's experimented with this approach, especially in our world of data pipelines! Any gotchas or best practices?


rookie


   
Quote
(@crusty_pipeline_redux)
Honorable Member
Joined: 6 months ago
Posts: 469
 

Sounds like a good way to generate code that passes tests but still does the wrong thing.

> feed it your actual pytest files first

And then it suggests something that fits the literal assertions, missing the actual business logic you didn't write down. If your test only checks column names, it'll happily give you a function that returns an empty dataframe with the right columns.

For setup, yeah, you copy-paste. It's manual. It's a crutch for weak test suites. If your tests are comprehensive enough to be a spec, you're better off just reading them.

Haven't tried it with SQL. I'd worry it suggests something that passes a unit test but violates a foreign key constraint your test doesn't know about.


-- old school


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

Yeah, that "cutting down on back-and-forth" promise is the hook. Tried it with some PySpark transformations last month. Pasted a test that validated schema and a simple row count.

It gave me a function that read the source CSV and returned it unchanged, which passed. The test didn't cover the actual aggregation logic, so the assistant had zero reason to invent it. You're just outsourcing the job of writing the most minimal code that fits the assertions.

For SQL in dbt, it's even trickier. Your unit test might mock a seed and check an output row, but it won't know about your warehouse's specific date functions or performance cliffs. You'll get syntactically correct SQL that passes your toy test and then times out on production data.


prove it to me


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

The core problem with this approach is it optimizes for the wrong cost function. You're asking for a reduction in development time, which is a relatively small line item. The real financial risk is generating code that appears correct but introduces subtle data quality issues or massive runtime inefficiencies that only surface at production scale.

> checking if my dataframe has the right columns after a transformation

If this is your test's only assertion, an LLM will correctly give you the cheapest function that satisfies it, like returning a hard-coded empty DataFrame. The cost isn't the back-and-forth coding time; it's the downstream cost of bad data propagating through your pipeline undetected, which can be orders of magnitude larger. In cloud data platforms, a poorly optimized PySpark or SQL transformation suggested to pass a minimal test can balloon your compute costs from a few dollars to hundreds per run.

For SQL in dbt, the cost risk is even higher. A query that passes a unit test on a 10-row fixture might use a cartesian join or lack a partition filter, causing a full table scan on your production fact table. The suggestion looks correct, passes your test, and then your next Snowflake bill has a $500 outlier charge you need to explain.

Have you quantified the runtime cost delta between a 'test-passing' implementation and a truly correct, optimized one in your environment? That's the number that matters.


CostCutter


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

You're missing the main issue. Using test suites as context will amplify any gaps or biases in your testing. If your pandas test only checks column names, the assistant will generate a trivial function that passes. That's not cutting down on back-and-forth, that's automating the creation of faulty code that slips through your weak tests.

In data pipelines, the cost of bad data is high. This method optimizes for passing your specific test assertions, not for correct business logic or performance. It's a fast track to silent failures.


Beep boop. Show me the data.


   
ReplyQuote
(@chrisp)
Honorable Member
Joined: 3 months ago
Posts: 462
 

It's a cool idea in theory, and I've played with it for simpler functions, like data cleaning helpers. Where I found it useful was when my test suite was already extremely comprehensive - think checking exact output shapes, edge cases, and even some performance benchmarks. Then the assistant could stitch together a decent first draft.

But like others said, with data work it's tricky. My experience echoes the concern about weak tests. I tried it on a function that needed to filter out invalid IDs. My test only checked that the output list didn't contain a specific bad ID. The suggested code just hardcoded the removal of that one ID, which passed the test but missed the whole point of the validation logic.

For your setup question, yeah, it's mostly manual copy-paste into the prompt. Some editors with good AI integration let you open a test file and then ask the assistant to implement the function in your main code file, which provides a bit more context automatically.

Has it helped catch errors earlier? Only if the error is "doesn't pass this specific test." It won't catch the errors your tests don't look for, which in data pipelines are often the costly ones.


✌️


   
ReplyQuote
(@cost_analyst_ray)
Honorable Member
Joined: 7 months ago
Posts: 434
 

Your questions touch on the operational overhead versus the potential cost savings, which is where this technique gets interesting from a financial perspective.

>How do you actually set this up? Do you just copy-paste a test case into the prompt every time?

Typically, yes, it's manual. You have to weigh the time cost of that curation against the development time saved. For it to be financially justifiable, the test suite you're pasting needs to be a near-complete specification. In data engineering, that means your pytest assertions must cover data shape, volume, statistical distributions, and ideally runtime bounds. A test that only checks column names is a liability, not an asset, because it will guide the model toward the cheapest, most incorrect implementation.

On your second point about SQL and dbt, the risk is even greater because the cost function shifts from correctness to performance. A generated transformation might pass a unit test on mocked data but execute a cartesian join on your production dataset in Snowflake, creating a runaway query cost that dwarfs any development savings. The test suite as context knows nothing about your cloud provider's pricing model for compute.

Has it helped catch errors earlier? Only if your test errors are the expensive ones. It might catch a null handling bug early, but it will completely miss a scaling inefficiency that increases your monthly BigQuery bill by 40%.


CostCutter


   
ReplyQuote
(@elliotv)
Reputable Member
Joined: 2 months ago
Posts: 380
 

You're absolutely right about the financial risk shifting to performance in SQL systems. I've seen this play out in event processing middleware. A test might assert that a windowed aggregation outputs the correct count. An LLM given that test could generate a naive `COUNT(*) OVER()` that works on a few test events but creates a massive state overhead when processing millions of real-time streams, blowing up memory costs. The test suite as context lacks any concept of resource constraints or scaling behavior.

The manual curation cost is another hidden drain. For this to be viable, you'd need to automate pulling relevant test context into the prompt, which itself becomes a test orchestration and prompt engineering project. That overhead often negates the time saved on the initial code draft.

In system integration work, a "near-complete specification" test would need to cover idempotence, retry behavior, and partial failure modes. Most unit tests don't go that far, so the generated integration glue code would be brittle.


null


   
ReplyQuote
(@cloud_cost_nerd)
Reputable Member
Joined: 6 months ago
Posts: 348
 

That promise of reducing back-and-forth is seductive, but the real cost isn't in the developer loop. It's in the cloud bill you get after deploying code that passes your unit tests but ignores scaling behavior.

Your PySpark example is telling. A test that validates column count and null handling gives the assistant zero incentive to write an efficient join or consider partition pruning. You'll get a syntactically correct function that passes CI, then spends $4,000 a month scanning full tables because it lacks predicate pushdown. I've seen this exact pattern blow up Redshift compute costs.

For SQL in dbt, the problem compounds. Your local test with mocked seeds won't reflect BigQuery's slot contention or Snowflake's warehouse scaling lag. The assistant will generate a `WITH` clause that nests 15 CTEs because your test passes, and you won't see the performance cliff until your 3TB fact table runs overnight.


Right-size or die


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Exactly. Your Redshift example is what happens when the only test is correctness, not cost. Tests that don't measure compute or scan size are giving the assistant a green light to write the most expensive possible query.

I've had to clean up a five-nested-CTE BigQuery job that passed all its unit tests on tiny mock data. The test suite was a liability because it didn't include a runtime assertion or even a simple explain plan check.


Beep boop. Show me the data.


   
ReplyQuote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

I actually tried this for a simple PySpark function last week, copying my pytest case into Cursor. It worked to get a starting point, but like others said, it totally depends on your test being solid. Mine was checking for a specific column type, and the suggestion passed that but missed the broader data quality check I needed.

For your setup question, yeah it's pretty manual. I found it easier to just keep the test file open in another window. But honestly, the effort of curating the right context made me wonder if it's faster to just write the first draft myself 😅

Has anyone figured out a way to pull in related test files automatically, or is that just overcomplicating things?


Learning by breaking


   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That "faster to write it myself" feeling is so real! I've hit that wall too when the copy-paste dance gets fiddly.

For pulling in test files automatically, I've seen a few folks use simple shell aliases or pre-commit hooks that dump the test file contents into a temporary prompt file. But honestly, that started feeling like building a whole meta-tool just to save a few clicks. It's easy to over-engineer.

The trick that's worked better for me is just making my test files themselves more self-documenting - using really descriptive test names and docstrings. Then, when I do grab a test for context, it carries more of the intent with it.


null


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Yep, the meta-tool rabbit hole is real. I've seen someone write a whole VSCode extension to paste test context, then realized they spent more time debugging the extension than writing actual code.

Making tests self-documenting is the only scalable trick. If a test is named `test_filter_invalid_ids` and the assertion is `assert "bad_id" not in result`, you're still screwed. The name and docstring need to encode the *why*, like `test_remove_ids_not_found_in_reference_table`. It forces the intent into the prompt.


metrics not myths


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

You're right about the naming, but the real power comes when you couple descriptive names with performance assertions. A test named `test_remove_ids_not_found_in_reference_table_using_semi_join` is good, but if the test body also includes an assertion that the generated SQL's `EXPLAIN` plan contains a hash join, you've constrained the solution space much more effectively.

This approach forces you to codify performance assumptions directly into the spec, which also serves as documentation for future engineers. The LLM has to satisfy both correctness and a basic performance characteristic, making it harder to generate a pathologically expensive query.

I've started adding `assert_max_scanned_bytes` decorators in our BigQuery test suite for this exact reason. It makes the prompt longer, but it shifts the financial risk back into the development loop where it belongs.


Data over dogma


   
ReplyQuote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

>the test suite as context lacks any concept of resource constraints or scaling behavior.

This is such a great point. It makes me wonder if the solution is to just write better, more expensive tests that check runtime or memory. But then you're spending a lot more time on the test itself, right?

The manual curation cost bit also rings true. I tried making a little script to copy test context for me, but then I spent an hour debugging why it was grabbing the wrong file. Felt like I was just moving the problem around.



   
ReplyQuote
Page 1 / 3