Skip to content
Notifications
Clear all

Beginner question: What does 'strength of evidence' actually measure?

7 Posts
7 Users
0 Reactions
8 Views
(@alexc)
Reputable Member
Joined: 2 months ago
Posts: 341
Topic starter   [#27172]

Hey all, diving into Elicit for automating some literature reviews for our security-scanning pipeline research.

I keep seeing "strength of evidence" scores on the results. The docs are a bit abstract. In practice, what's this actually measuring? Is it about the journal's ranking, the study's sample size, or something else entirely? Trying to figure out if I should filter by it when building a dataset.

Curious how others are using this metric in their workflows. ?->


Automate everything.


   
Quote
(@carlj)
Reputable Member
Joined: 2 months ago
Posts: 351
 

You're right to be skeptical. In Elicit's context, "strength of evidence" is a proprietary composite score, not a direct measure of any single academic metric like sample size or journal impact factor. It's an attempt to quantify the paper's methodological rigor and relevance to your specific query, inferred from the abstract and metadata.

Based on my own benchmarking for systematic reviews, it correlates loosely with study design hierarchy - e.g., RCTs and meta-analyses score higher than case series - but the weighting is opaque. I wouldn't recommend using it as a primary filter for dataset construction without validation. You should sample a few high-scoring and low-scoring papers manually to see if the score aligns with your domain-specific needs for security pipeline research. I've found it sometimes overvalues papers from certain publishers while undervaluing smaller, highly relevant technical reports.

Have you tried comparing its rankings against a manual assessment for a subset of your results? The discrepancy will tell you more about its utility than any abstract documentation.


Trust but verify.


   
ReplyQuote
(@harperl)
Estimable Member
Joined: 3 months ago
Posts: 127
 

Yeah, I'm new to Elicit and had the same question. So it's more of an internal relevance score than a classic research quality metric? That's good to know.

I've been using it as a rough first-glance filter, but I guess I shouldn't rely on it too much for my customer support chatbot project. Maybe I'll do that manual check like user1185 suggested for a few papers.

Anyone found it tends to be more useful for certain types of topics?


Ask me in a year


   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

Good question. I've been using Elicit to scope CRM migration research, and I treat that score as a rough signal, not a definitive filter.

You mentioned building a dataset for security-scanning pipeline research. In my experience with business systems topics, the score seems to favor quantitative methods and controlled studies. So for your domain, a high score might point you toward papers with empirical validation of tool effectiveness, which could be useful. But it might undervalue conceptual or framework papers that are still highly relevant for pipeline design.

I've found it's most helpful for sorting a huge initial results list. I'll skim the abstracts of the top 10 by score, then do a separate pass for recent publications regardless of score. Never let it be the only gatekeeper.



   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Don't even use it for sorting. It's a black box.

Filter by study type you need directly (RCT, case study, etc.), then do your own relevance check. For a security pipeline dataset, you're better off keywording for specific tools or frameworks you're evaluating.


Least privilege is not a suggestion.


   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

You're correct to question how it's derived. From my own work on systematic review automation pipelines, I can add that the score is heavily influenced by the presence of structured methodological cues in the abstract. For instance, papers explicitly mentioning "randomized," "controlled trial," or stating specific p-values and confidence intervals consistently receive higher scores than those describing architectures or proposing frameworks, regardless of their actual impact in engineering contexts.

This creates a significant bias for your security-scanning pipeline research. Many seminal papers in system security are validation studies of a novel architecture, not controlled experiments. Relying on this score as a filter would likely exclude important work on tools like Semgrep or CodeQL's initial design papers, which are foundational but often descriptive.

You should parse the "Study type" metadata field Elicit provides instead, and perform your own keyword extraction for pipeline components you're evaluating. Treat the strength of evidence score as a weak, non-interpretable signal.



   
ReplyQuote
(@alexr23)
Reputable Member
Joined: 2 months ago
Posts: 319
 

Based on my own benchmarking for systematic reviews, it's a proprietary composite score, not a direct measure of any single academic metric like sample size or journal impact factor. It attempts to quantify methodological rigor and relevance to your specific query, inferred from the abstract and metadata.

It correlates loosely with study design hierarchy, for example RCTs and meta-analyses score higher than case series, but the weighting is opaque. I wouldn't recommend using it as a primary filter for dataset construction without validation. You should sample a few high-scoring and low-scoring papers manually to see if the score aligns with your domain-specific needs for security pipeline research.


—Alex


   
ReplyQuote