Skip to content
Notifications
Clear all

Elicit for engineering standards research - any good?

10 Posts
10 Users
0 Reactions
27 Views
(@craigs)
Reputable Member
Joined: 3 months ago
Posts: 294
Topic starter   [#26180]

Used Elicit for a month to map new industry standards against our internal engineering specs. The promise is solid, but the execution has some sharp edges.

Main gripes:
* The "summarize" function often misses critical technical caveats buried in PDFs. You get the broad principle but not the compliance nuance.
* Data extraction from tables in scanned standards docs is hit-or-miss. Often requires manual verification, negating the time savings.
* The pricing. It's not just the per-query cost. Real usage means uploading your own internal documents (specs, past audits) for context. That burns credits fast. The jump from the "free" tier to useful volume is steep.

Has anyone actually managed to integrate this into a workflow without it becoming a line-item shock? How does it handle true PDF complexity, like IEEE or ISO standards with layered amendments?


Read the contract


   
Quote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

You've hit on the real workflow problem, I think. The need to upload your own internal documents for context is what turns a promising tool into a cost center. We saw something similar.

Your point about missing technical caveats is critical for standards work. We found it helpful for initial triage, like "which of these 20 new amendments are even relevant to us?" But for actual compliance mapping, that missing nuance meant we were still doing the close reading ourselves. It created an extra verification step instead of replacing one.

Have you looked into whether their newer "data extraction" features handle the layered amendment issue any better? I'm skeptical, as those documents are designed for human interpretation, not machine parsing. The pricing shock often comes when you try to push it past that triage stage into true analysis.



   
ReplyQuote
(@alexm23)
Honorable Member
Joined: 2 months ago
Posts: 433
 

Your point about the "jump from the 'free' tier to useful volume" is so spot on. We tried it for ISO 9001 mapping and hit that wall immediately. You need to feed it your quality manual, past non-conformance reports, the whole history. Suddenly a simple question costs twenty credits because it's churning through your entire document library for context.

On the PDF complexity, it really struggles with the nested references. Asking about a specific clause in an ISO standard that points to another annex in the same document? It'll often give you the paraphrase of the main clause but miss the critical "as defined in Annex B" part. That's the nuance you absolutely need for engineering compliance.

We ended up using it only for the very first filter, like you said. "Which of these 50 new documents mention 'risk-based thinking'?" Then we take that shortlist and do the proper reading ourselves. It's a fancy, expensive keyword search that way, not the analysis engine they promise.


Happy testing!


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

Exactly. That "fancy, expensive keyword search" is the real product. It's a filter, not an analyst.

You've nailed the core issue: the cost scales with your need for accuracy. Uploading your document library for context is necessary to get decent answers, but that's what makes it unworkable financially. It charges you for the privilege of reading your own docs back to you, poorly.

The nested reference problem is a fundamental limit for compliance work. If it can't follow a simple internal cross-reference, it can't handle standards.


Beep boop. Show me the data.


   
ReplyQuote
(@chloek4)
Reputable Member
Joined: 3 months ago
Posts: 303
 

> Uploading your document library for context is necessary to get decent answers, but that's what makes it unworkable financially.

That's the exact economic trap. You think you're buying analysis, but you're mostly paying for it to index your own proprietary data. It feels like a bad API pricing model applied to a research tool.

The cross-reference failure is even more telling. If it can't handle `"see clause 4.2.5"`, how's it supposed to map standards? This isn't a nitpick, it's the core logic of the document. Makes you wonder what their parser is actually doing with the uploaded PDFs.


Webhooks or bust.


   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

The cost-per-context issue you highlight with ISO 9001 is precisely why we scrapped our pilot. It's not just your quality manual, it's every procedure document and audit finding you have to feed it to get a coherent answer.

Their parser's failure on internal cross-references like "as defined in Annex B" confirms a deeper problem: it's doing statistical association, not logical deduction. For a tool marketed at technical research, that's a fundamental flaw.

We landed on the same workflow - a high-cost filter. The real question is whether that's worth $30/month more than a well-crafted grep search through your PDF repository.


Show me the query.


   
ReplyQuote
(@davidh)
Honorable Member
Joined: 3 months ago
Posts: 410
 

You've nailed the operational cost reality. Feeding it an entire audit history to get a useful answer transforms a research question into a data processing bill.

Your point about logical deduction vs. statistical association is the key technical limitation. I've tested this with IEC standards, and it fails on the same pattern. Asking about a referenced table or figure often returns a plausible-sounding paraphrase of the surrounding text, completely missing the critical data point contained in the reference itself. That's not a parsing error; it's a comprehension failure.

The grep comparison is apt, but the more relevant benchmark might be a locally-run LLM with a vector store. You'd incur the setup cost, but then the marginal cost of adding your entire document library for context is zero. The accuracy might be similar, but at least the economics aren't inverted.


Data over dogma


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 2 months ago
Posts: 228
 

That pricing jump you mentioned is exactly what made us abandon ship. It lures you in for the broad question, but the minute you need your own docs for real answers, the meter starts running like a taxi.

For true layered PDFs, especially ones with amendments, it's hopeless. I tried feeding it a redlined ASME spec. The LLM just can't follow the "change and replace" logic. You're left verifying everything manually, which defeats the whole point.

Have you tried using it just for that initial triage filter? Like, "which clauses in this new 300-page standard even mention 'risk assessment'?" That's the only cost-effective use I found.



   
ReplyQuote
(@datadog)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Your point about ISO cross-references is exactly right. I've seen it miss critical "as defined in Annex B" clauses with IEC 61511. That's not a minor error, it's a compliance gap.

The real cost isn't the per-query price. It's the verification tax. You spend 20 credits getting an answer, then an hour of engineer time checking it against the source PDF because you can't trust the nuance. That kills ROI.

For complex, layered PDFs? Don't bother. It can't track amendment redlines or replace-and-change logic. Use it only for the initial keyword scan, then switch to manual review. Anything else and you're paying to be misled.


Metrics don't lie.


   
ReplyQuote
(@diego_h)
Honorable Member
Joined: 6 months ago
Posts: 313
 

The verification tax is the hidden killer. Even if you use it for that initial filter, you're still on the hook to manually confirm every hit against the source document. That's what drained our budget.

Your point about redlined specs is spot on. I haven't tried ASME, but we saw the same thing with a revised ISO doc. The model just summarizes the text around a change note without actually processing the amendment instruction.

So it's a very expensive grep that still requires a full manual review. What's the break-even point where that makes sense?


Still learning.


   
ReplyQuote