Skip to content
Notifications
Clear all

TIL: You can use custom keywords to bias Scholarcy's extraction

4 Posts
4 Users
0 Reactions
19 Views
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
Topic starter   [#26105]

While evaluating Scholarcy's performance for a technical literature review on eBPF-based observability tools, I noticed its default extraction was prioritizing generic "conclusions" over the specific metrics and architectural patterns central to my research. This led me to investigate its configuration options, where I discovered the under-documented but powerful "Custom Keywords" feature. This feature allows you to bias the algorithm's sentence scoring, significantly improving the relevance of the extracted summary for specialized domains.

The mechanism is straightforward. Scholarcy assigns a score to each sentence in a document based on various linguistic and positional features. By providing a list of custom keywords, you artificially inflate the score of sentences containing those terms, making them more likely to be included in the summary flashcard. This is particularly useful for technical papers, RFCs, or documentation where domain-specific terminology is crucial.

You can access this feature via the "Library" dashboard. The process is as follows:
1. Navigate to your "Library" and locate the article.
2. Click on the three-dot menu (`...`) next to the article title and select "Edit settings."
3. In the pop-up modal, you will find a field labeled **"Custom keywords (separated by commas)"**.

Here is an example configuration I used for a paper on Kubernetes autoscaling:
```
horizontal pod autoscaler, HPA, custom metrics, PrometheusAdapter, v2beta2, utilization threshold, readiness probe, scaling latency
```

The impact on the resulting summary was substantial. Without custom keywords, the top flashcard sentences discussed general challenges of cloud scalability. With the keywords applied, the extraction immediately prioritized:
* The specific API version (`autoscaling/v2beta2`) enabling external metrics.
* The architecture diagram description linking Prometheus metrics to the HorizontalPodAutoscaler via the custom metrics API.
* The numerical threshold values for CPU and memory-based scaling.

For reproducible research or systematic reviews, this feature can be standardized. I maintain a set of keyword lists for different subdomains (e.g., `cost_optimization_keywords.txt`, `service_mesh_keywords.txt`) and apply them consistently to batches of imported papers. This transforms Scholarcy from a general-purpose summarizer into a tailored research assistant that understands the lexicon of your field.

Potential pitfalls to consider:
* **Over-bias**: An excessively long keyword list can overwhelm the native scoring, potentially pulling in low-significance sentences that merely contain terms.
* **Synonymy**: The matching appears to be exact. Using "latency" may not capture "response time" unless both are listed.
* **Document Structure**: The feature biases sentence selection but does not fundamentally alter the underlying extraction model. Poorly structured source documents may still yield suboptimal results.

In conclusion, the custom keywords feature is a powerful lever for technical users. It provides a necessary layer of control, aligning the tool's output with your specific information retrieval goals. For those of us working in dense, jargon-heavy fields, it dramatically increases the utility of automated summarization.



   
Quote
(@brianl)
Honorable Member
Joined: 3 months ago
Posts: 506
 

That's an interesting find about the custom keywords. I've been using Scholarcy for summarizing some of the vendor API documentation and whitepapers we get from our logistics software partners, and I've run into the same issue where it latches onto high-level business benefits instead of the actual technical implementation steps.

> The mechanism is straightforward.

This part made me wonder about the potential for over-bias. In your testing with the eBPF materials, did you find there was a point where adding too many keywords started to degrade the quality of the summary, maybe by pulling in sentences that were just term definitions without the surrounding context? I'm trying to gauge how selective I need to be with my own keyword lists for supply chain protocols.



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

I appreciate the detail in the post, particularly the workflow you outlined for accessing the feature via the Library dashboard. I've been using this on infrastructure RFCs and security compliance docs for about six months, and you've nailed the core mechanism.

One caveat you might want to add: the keyword scoring is additive. If you only bias the algorithm without any curation, you'll end up with a disjointed list of sentences that just happen to contain your terms. It doesn't magically understand context or relationships. You still need to review the output, and sometimes it's better to use a tighter list of 5-7 highly specific keywords or key phrases rather than 20 individual terms. For example, with eBPF, I'd use "kprobe," "user probe," "map types," and "tail call" rather than just "eBPF," "observability," and "performance."

Have you looked at whether this biasing works better on the full-text extraction versus the abstract? I've found it's far more effective once you're past the introductory sections of a paper.



   
ReplyQuote
(@chrisf)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Oh, I didn't know you could do that. That's really helpful for cutting through the fluff in technical docs.

>via the "Library" dashboard
Is that feature available on the free plan? I've been using the browser extension and haven't seen a settings menu there.


Still learning.


   
ReplyQuote