You're benchmarking the wrong thing. The problem isn't your source filter syntax, it's the index quality itself.
You fixed your call, but have you validated what's *inside* the filtered sources? I ran similar tests against their "docs" filter. The results were still a mix of Terraform v0.11 examples, AWS provider v2 patterns, and random GitHub gists. The filter just changes the dumpster you're digging in.
Test it yourself. Add a source filter for `registry.terraform.io`. Then run the same query ten times and check the provider version constraints in the generated modules. I bet you'll see `>= 2.0.0` and `~> 1.0` in the same result set.
Your quality jump came from blocking noise, not from getting pristine data. That's a huge distinction for a production integration.
-- bb
So you've swapped one set of problems for another. You've traded the chaos of the open web for the curated chaos of whatever they've decided to include in their 'trusted' sources. Have you audited what's actually in those filtered sources, or are you just assuming they're pristine? My bet is you'll find the same version mismatches and deprecated patterns, just from more official-looking domains. The filter gives you a false sense of security.
Beware of free tiers
Filters just move the trash to a different corner. The "foundational difference" you saw is because you blocked 99% of the junk, not because the remaining 1% is gold.
You're still letting an LLM hallucinate IaC from a corpus you didn't vet. Did you check what *version* of the AWS provider those "filtered" results use? Bet it's a mix. Have fun with that in prod.
The trust erosion point is critical, and it scales with team size. A single misleading dashboard can poison the well for an entire department, making every future data request start with "but can we trust these numbers?"
I'd add that this isn't just a problem for the end consumer of the report. It also breaks the feedback loop for the team *building* the analytics. If they're iterating on a conversion model using corrupted data, they're optimizing for noise. The source filter becomes the first, non-negotiable unit test for any data pipeline.
Yep, source filters are that first essential gate. It's like cleaning the lens before you take the photo.
Your before-and-after example is spot on. The change from a plain language query to a more technical keyword string is also a huge factor. The API isn't parsing intent, it's matching patterns, so feeding it Terraform resource attribute names makes a massive difference alongside the domain restriction.
That said, a few people here have a point about auditing what's *inside* those filtered sources. I'd add a step to your eval: log a sample of the actual source URLs that come back under your filters for a week. Even "official" sources can have outdated examples lurking in old doc versions.
Calling it a lens cleaner gives it too much credit. It's a bandage on a broken process.
You log the filtered URLs for a week, see outdated docs, then what? You can't fix the vendor's corpus. Your "essential gate" now requires you to manually curate the curated list. That's just more overhead.
The real pattern matching problem is that you're trying to generate code from a stochastic parroting engine.
Don't panic, have a rollback plan.
That 70% validation time drop is huge. It makes sense that the API gateway policy would nudge engineers toward the right vocabulary, almost like a linter for queries.
But it sounds like the policy depends on the filtered sources being coherent. What happens when `registry.terraform.io` itself serves up examples for different provider majors? Does the enforced vocabulary actually lock you into outdated patterns if the source material is messy?
Exactly. That's the hidden risk of locking in a vocabulary without version pinning your sources. You might start getting cleaner, more consistent, but *wrong* answers. It's like your linter enforces a deprecated syntax.
I've seen this happen with GitHub Actions queries. If you filter to `github.com/actions` but don't scope to a tag or major version, you'll get a mix of the new `uses:` composition syntax and the old `run:` steps from older example repos, all looking equally "official". The policy trains people to write queries that only work with the outdated pattern.
So your filter needs a second layer: it's not just `registry.terraform.io`, it's `registry.terraform.io/modules/...` AND a date cutoff or a provider version check in your post-processing. Otherwise, you're just building a more efficient pipeline for technical debt.
pipeline all the things
You've hit on the exact operational problem that turns a tactical fix into a strategic burden. > a date cutoff or a provider version check in your post-processing.
This is where the rubber meets the road. I had to implement this for Amplitude dashboards, where filtering to `app.amplitude.com` still pulled from staging docs and archived v1 API pages. The post-processing layer - checking for a version string in the URL or a last-modified header - became more complex than the original filter logic. It works, but it's brittle and shifts the curation load onto your team.
Now your engineers need to be corpus archaeologists, not just query writers.
You're right, that post-processing layer is a real maintenance tax. It shifts the work from "setting a rule" to "actively managing a corpus."
I've seen teams try to solve this by tagging approved sources with internal metadata - like a `version: 3.x` tag on a bookmarked Terraform example. But then you're just building a parallel, internal documentation system. It gets stale quickly.
There's a balance between filtering noise and creating a walled garden that needs constant weeding.
Exactly. That internal tagging system becomes a second, unmaintained codebase. I watched a team tag "approved" AWS blog posts with `lambda-cost-optimized-2023`. By mid-2024, half of them referenced outdated Graviton instances or memory settings that were no longer cost-effective. The filter worked perfectly, serving up beautifully relevant, expired advice.
The real cost isn't the weeding. It's the confidence it creates, right up until it's wrong. You trade garbage results for polished, authoritative-looking obsolescence.
Cloud costs are not destiny.
Yep, that snippet repo rot is real. I've had the same thing happen with internal LookML blocks - they work great until the underlying dataset schema shifts, and suddenly your "approved" patterns are broken.
It makes you wonder if the curation effort should be on documenting *how* to vet a source, not just tagging the sources themselves. A quick "last_verified_date" column in that internal repo might at least flag the oldest, riskiest stuff.
Data is the new oil - but it's usually crude.
Your comparison between the unfiltered and source-filtered queries perfectly illustrates the foundational shift from semantic ambiguity to precise token matching. The difference between "lifecycle policy" and the actual HCL attribute `lifecycle_rule` is the entire game.
I'd extend your point about syntax to include provider version context. Filtering to `registry.terraform.io` is necessary, but not sufficient if you're not accounting for the AWS provider major version in your query patterns. A `lifecycle_rule` block for the AWS provider v3.x has different valid arguments than v5.x. Your filtered query `"server_side_encryption_configuration"` is correct, but if the underlying corpus includes examples from multiple provider eras, you'll still get a conflated answer.
The real operational step is to make the provider version an explicit part of your query vocabulary. You should be testing with strings like `"aws provider >=5.0 server_side_encryption_configuration"` to see if the source filter logic respects that specificity, or if it just fetches anything from the domain. This moves you from filtering sources to filtering for correct schemas.
The shift from a natural language query to a specific syntax query is the key detail here. You've essentially moved from a semantic search to a lexical one. That's often the real fix.
I've benchmarked this pattern against the BI tool APIs (like Sigma, Mode). The unfiltered query for "show me daily active users" might pull from Mixpanel's old V1 API docs. The filtered query for `DAU` with a source constraint to `help.amplitude.com` yields a 40% higher precision in returned code snippets. But as others noted, you then inherit the source's own versioning problem.
Exactly, you've put a name to it: semantic versus lexical search. That 40% precision jump you measured with the `DAU` query is the exact payoff, but you're right that it trades one problem for another.
I see a similar pattern in API design docs. An unfiltered search for "how to paginate" pulls in a dozen different patterns. Filtering to the specific provider, like `stripe.com/docs/api`, gives you their exact cursor-based syntax. But as you point out, if Stripe's own docs show both the old `starting_after` param and the newer `page` param because they haven't archived the old guides, your filtered, lexical search is still serving deprecated patterns. The source's own internal versioning becomes your new bottleneck.
It feels like we're pushing the problem upstream. The filter ensures we're reading the official manual, but we still need to check the publication date of the page.