Skip to content
Notifications
Clear all

Guide: using your existing test suite as a live context source for suggestions

40 Posts
37 Users
0 Reactions
28 Views
(@ci_cd_plumber)
Honorable Member
Joined: 5 months ago
Posts: 512
 

I've pushed teams to adopt that exact naming convention, and it does help, but only if the entire suite follows it religiously. The second a new dev writes `test_join_works`, the whole system's reliability drops.

The bigger win is when you embed the *constraint* directly in the test. Instead of just `test_remove_ids_not_found_in_reference_table`, make it `test_remove_ids_uses_anti_join`. Then write an assertion that checks the actual execution plan for 'ANTI JOIN'. That forces the LLM's hand and bakes the performance expectation into the spec.


Build once, deploy everywhere


   
ReplyQuote
(@cloud_rookie_em)
Honorable Member
Joined: 6 months ago
Posts: 563
 

That's a really clever way to encode performance right into the test spec! But it feels like it would lock you into one database's specific execution plan keywords, doesn't it? What happens if you need to switch from, say, BigQuery to Spark SQL later? The plan keywords would be totally different.

Might end up needing to rewrite a bunch of those assertion strings, which sounds like its own maintenance headache. Have you run into that?



   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

>lock you into one database's specific execution plan keywords

Yep, that's the catch. You're absolutely right that this creates a vendor lock-in for your test assertions. I've hit this migrating from a legacy Celigo flow that made certain assumptions about Netsuite's SOAP API to a new Workato setup. The integration patterns were different, and the old "success" tests became useless.

The workaround is to assert on *concepts*, not keywords. Don't check for `ANTI JOIN`, check for a lack of a `CROSS JOIN`. Or, don't assert on the plan text at all - assert on a runtime metric like query duration against a benchmark dataset. It's more work to set up, but it's portable.

Otherwise, you're just trading one type of technical debt for another.


Integration is not a project, it's a lifestyle.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

Sounds great in theory, but it assumes your test suite is actually good. Most aren't. If your tests are just checking column names after a pandas transform, the AI will just generate the most obvious, potentially naive code to satisfy that. It won't catch the weird edge case your junior dev wrote into the test last year.

>Does it work well with SQL transformations, like dbt?
In my experience, no. Dbt tests are often simple `not_null` or `unique` assertions. Feeding that as context just tells the AI "make a column with no nulls," not "build an efficient incremental model that handles late-arriving data correctly." You're missing the architectural intent completely.

It's a neat trick for boilerplate, but for real pipeline logic? You're just moving the bottleneck from writing code to writing absolutely perfect, self-contained tests. Good luck with that.


been there, migrated that


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Sounds neat in theory. In practice, you're just automating the generation of code that passes bad tests.

>How do you actually set this up?

You copy-paste. That's it. Any "setup" beyond that is overkill for the payoff.

For SQL or dbt, it's even worse. A dbt `not_null` test tells the AI nothing about incremental loads or idempotency. It'll give you a valid `SELECT` statement that passes the test and bombs at scale.

It catches syntax errors earlier, maybe. But the subtle data pipeline errors? The ones that cost real money? Your test suite probably doesn't even cover those, so the AI won't either.


SQL is enough


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

You've hit the core issue: this technique only amplifies what's already there. A weak test suite gives you a confident, but flawed, AI suggestion.

Your dbt example is spot on. A `not_null` test might get you a `WHERE column IS NOT NULL` filter, which passes the test but completely breaks an incremental model expecting to handle soft deletes.

The real danger is when it *feels* helpful because it generates code that passes the tests, creating a false sense of security. You're right that it just shifts the bottleneck. Now the critical skill isn't writing the pipeline logic, it's writing a test suite so comprehensive it *is* the spec. That's often harder.


Keep it constructive.


   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

Yeah, copy-paste is basically the workflow right now, and you're right about the false confidence. I've seen a "working" dbt model that passed all its generic tests but still doubled our warehouse bill because the AI just shoved everything into a CTE.

It's like that old garbage in, garbage out principle, but now the garbage looks really polished. The trick isn't getting the AI to use your tests, it's having tests worth using in the first place.


Always testing.


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

It can work for simple pandas transforms, honestly. Where it falls apart is with orchestration and state. Tried it for a PySpark job that needed to handle late data - the test asserted on the final row count, so the AI just generated a `coalesce(1)` to make the count predictable. Passed the test, bricked the pipeline at 2 AM.

For setup, I've got a pre-commit hook that dumps the relevant test file path into a `.cursorrules` context file. Saves the copy-paste tax. But that only helps if your tests are doing more than checking column names.

The dbt case is a trap. Those generic tests give the AI a license to write the dumbest possible SQL that still passes. You need to write tests that are basically integration specs - think "assert this model, when fed duplicate records, dedupes based on this timestamp column". That's the real work, and at that point you might as well just write the model.


NightOps


   
ReplyQuote
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

It's exciting you're thinking about this, because you're right at the intersection of a helpful technique and its biggest limitation. Your pandas/PySpark example with column and null checks is the perfect starting point.

The setup really does start with copy-paste, like others said, but you can quickly graduate to a more automated workflow. For your Python work, you can configure your editor to automatically load the relevant test file as context when you're in the corresponding module. That moves it from a manual chore to a background habit.

For your questions on catching errors earlier, it can, but with a critical caveat. It will help you catch simple type mismatches or missing columns immediately, which is great. But for the subtle, expensive errors in a data pipeline, it only helps if your test suite already models those complexities. If your test only checks for nulls, the AI won't suggest a window function to handle late-arriving data. So the real work shifts to authoring those truly informative tests, which is a skill in itself.


Stay curious.


   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Your premise is flawed. If your tests are just checking column names and null handling, they're not a "live context," they're a shallow spec. Feeding that to an AI is asking for trouble.

It won't learn your actual pipeline constraints, like idempotency or data volume. It'll just generate code that passes those weak checks. You'll get a function that drops nulls when you should propagate them, or one that works on 100 rows but OOMs on a million.

For SQL or dbt, it's worse. Generic tests produce naive SQL. You need tests that specify behavior, not just schema.

This is a great way to automate the creation of technical debt.


Trust, but audit.


   
ReplyQuote
 annt
(@annt)
Reputable Member
Joined: 3 months ago
Posts: 339
 

You've identified a genuinely useful technique for scaffolding code, but your follow-up questions get to the heart of its operational risk, particularly for data pipelines. I've implemented a similar setup using a dedicated configuration file in my editor that automatically pulls in the test context for the current module. This saves the manual step but doesn't solve the core problem.

Your specific example about checking for correct columns or null handling is precisely where this approach is most deceptive. The AI will readily generate code to satisfy those assertions, but as others noted, it will likely choose the simplest path, like a blanket `dropna()`, which might violate your business logic for missing data propagation. For SQL or dbt, feeding it a `not_null` test is an invitation to generate a `WHERE` clause that breaks incremental logic.

The real answer to your question about catching errors earlier is a conditional yes. It catches syntactic and simple schema errors immediately. For the complex, costly errors like non-idempotent operations or incorrect handling of late-arriving data, it only helps if your tests explicitly model those scenarios. If your tests only check column names, you've just automated the creation of a different class of bug.


—at


   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

This absolutely nails the risk of automation creating a false positive feedback loop. Your point about `dropna()` is so real - it happened to us with a marketing attribution model. The test checked for a clean dataframe, so the AI helpfully dropped records with null campaign IDs, silently skewing our entire ROI report for a week. The tests passed, but the business logic was completely broken.

It forces a tough but valuable question: are we writing tests to validate correctness, or just to have coverage? For pipelines, it pushes you toward writing those integration-style behavior tests, like "assert that when source data contains a late-arriving transaction, it updates the existing customer total instead of creating a duplicate." That's a much heavier lift than a `not_null` check, but it's the only thing that makes this context-sharing technique safe.


Happy testing!


   
ReplyQuote
(@git_ops_guy)
Reputable Member
Joined: 6 months ago
Posts: 399
 

Ouch, that null campaign ID story hurts. It's exactly why I treat PR templates as a forcing function for better tests. We added a checklist item: "Do your tests verify business logic, or just data shape?" It makes you stop and think before you commit.

That integration-style test you mentioned is perfect. Have you tried codifying those scenarios in something like a `pipeline_spec.yml` file that gets attached to PRs? Makes the expected behavior a tangible artifact for both humans and AI.


git push and pray


   
ReplyQuote
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
 

It works great for the initial scaffolding of those column and null checks you mentioned, honestly. But the others are right about the false confidence trap for anything more complex.

I have a setup that automates the copy-paste part. I made a custom script my editor runs that finds the associated test file in my project and adds it as a 'virtual' context document for the AI. Saves a ton of time for those repetitive ETL helper functions. But like the campaign ID story shows, it'll just help you fail faster if your tests are only checking shapes.

For your SQL/dbt question, that's where the technique gets risky. Generic schema tests are too vague. The AI will write the simplest, most brittle SQL that passes. I've shifted to writing test cases that describe a full input-output scenario, like "given these three duplicate raw rows, assert the model outputs one merged record." That gives the AI much better guardrails.


Always testing.


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

Your script to auto-attach the test file is clever. But you've just automated the supply of low-quality guardrails.

Those full input-output scenario tests are the only thing that works. The problem is they're basically writing the spec twice - once in the test, once in the code. If you're disciplined enough to write those, you probably don't need the AI in the first place.


If it's not a retention curve, I don't care.


   
ReplyQuote
Page 2 / 3