Skip to content
Notifications
Clear all

Profound vs LLM Pulse - which produces less hallucination in technical writing?

37 Posts
35 Users
0 Reactions
100 Views
(@graces)
Reputable Member
Joined: 3 months ago
Posts: 441
 

It's a helpful start for your evaluation, but I think you're right to look for the omissions. Your LLM Pulse summary suggests it included `on_schema_change`, which is a more advanced detail that indicates it might have drawn from slightly more complete training data.

However, that's still recall. To really test for hallucination, you'd need to see if either tool could correctly explain the *interaction* of those parameters in a non-standard scenario. For instance, asking when `unique_key` can be optional, or what happens if `on_schema_change: 'sync_all_columns'` is used with a strategy that doesn't support it.

A tool that's good at technical writing will often signal its own limitations, like adding "under the default `merge` strategy" when explaining `unique_key`. The absence of those qualifiers can be a subtle form of hallucination, implying a universal rule where one doesn't exist.


Stay curious.


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Exactly. That missing qualifier is the silent killer. I've been burned by a model that correctly stated "use `unique_key` for incremental models" without the critical "for the `merge` strategy" caveat. I spent half a day debugging why my `append` model was creating duplicates before I realized the advice was context-blind.

The real test is asking a question where the official documentation itself is ambiguous or has a known quirk. If the tool just regurgitates the doc's phrasing without integrating known community workarounds or platform-specific bugs, it's not reasoning, it's a fancy mirror. I'd rather have a tool say "the docs say X, but a common issue on BigQuery is Y" than give me a perfectly recalled but naively incomplete fact.


Pipeline is king.


   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

The BigQuery example you picked is spot on. It's not just about missing a qualifier, it's about missing the entire *ephemera* layer that exists outside official documentation.

I'd take it a step further. The most dangerous hallucinations happen when a tool *does* integrate a known community quirk, but gets the implementation details subtly wrong because its training data mixed a Stack Overflow answer from 2020 with a 2023 API change. You end up with a confidently presented, "enhanced" answer that's more wrong than a simple doc regurgitation would have been.

A true test for these tools would be to ask something like, "How do I handle incremental loads on Snowflake with a Type 2 slowly changing dimension?" Any answer that doesn't immediately caveat its response with "depends on whether you're using `dbt-core` or `dbt-snowflake` version >= 1.3" is fundamentally unreliable, because the available strategies changed there. It's that dependency-aware contextual layer that separates a useful assistant from a textbook.


Measure twice, cut once.


   
ReplyQuote
(@emilyh)
Estimable Member
Joined: 2 months ago
Posts: 166
 

That point about mixing old Stack Overflow answers with new API changes really resonates. I've been trying to automate some web scraping, and I'll find a tool's answer that perfectly combines a deprecated library method with a modern authentication flow. It looks coherent, but it's a dead end.

It makes me wonder, for a test like your Type 2 SCD question, how would you even verify the answer is right? If the tool correctly includes the dbt-snowflake version caveat, you'd still have to check that caveat is actually true. You're just trusting it recalled a different piece of documentation correctly. Isn't that just pushing the verification problem one step back?



   
ReplyQuote
(@emmaw)
Estimable Member
Joined: 3 months ago
Posts: 139
 

Thanks for sharing the test! This is really helpful.

I'm new to dbt myself, and just knowing about the `is_incremental()` macro is a big tip for me. It seems like a small detail, but it's the exact kind of practical bit that helps you actually build something.

Did either tool mention what happens if the incremental load fails partway through? Like, is your data in a weird state? That's my biggest fear when trying something new.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Exactly. The citation problem is the same one we hit when we started trying to formalize our internal playbooks. If you limit the ground truth to vendor docs, you bake in the vendor's own omissions and mistakes.

I ran into this last month with the Azure Terraform provider. The docs for a specific `azurerm_kubernetes_cluster` update block were wrong for a minor version. The correct behavior was only documented in a closed GitHub issue from six months prior. Any tool trained only on the official Azure schema would confidently give the wrong answer, citations and all.

Your benchmark idea is the key. You need a corpus of "tribal knowledge" prompts where the source is an issue tracker comment, a release note deep in a changelog, or a forum post that corrected a core misunderstanding. Then you measure if the tool can both find it and attribute it correctly, not just default to the pristine, outdated manual.


Automate everything. Twice.


   
ReplyQuote
(@danielr)
Reputable Member
Joined: 2 months ago
Posts: 408
 

> "tribal knowledge" prompts

Good idea in theory, but you're betting on the tool having clean access to that noise. Issue trackers and forums are a graveyard of deprecated fixes and angry rants. How does the model weight a correct 2023 GitHub comment against a popular but wrong 2021 Stack Overflow answer?

This isn't just a data problem. In procurement, we see this all the time: vendors let their docs rot because they know the community will patch the gaps for free. You're benchmarking against a corpus they outsourced to their own users.

And what about proprietary tools? The real tribal knowledge for many enterprise SaaS platforms lives in locked Slack channels or partner portals. Your benchmark fails the moment you step off the beaten path of open source hype.


Trust but verify.


   
ReplyQuote
Page 3 / 3