Skip to content
Notifications
Clear all

What retrieval pipeline actually works for a 100k document medical library

4 Posts
4 Users
0 Reactions
2 Views
(@emmae)
Reputable Member
Joined: 2 months ago
Posts: 255
Topic starter   [#29539]

Hi everyone! I've been tasked with helping our medical research team set up an internal knowledge base using Relevance AI. We have a library of around 100,000 documents—mostly PDFs of clinical studies, whitepapers, and regulatory guidelines. The goal is to have a chatbot that can accurately answer specific questions from this collection.

I've gone through the tutorials and set up a basic pipeline with chunking and embedding, but the answers I'm getting back are... not great. They either pull from completely unrelated documents or miss the key details. I think my chunking strategy might be wrong for such dense, technical material.

What retrieval pipeline configuration has actually worked for you at this scale, especially with complex text? I'm particularly unsure about:
* Chunk size and overlap for long, detail-heavy PDFs.
* Whether to use their hybrid search (keyword + vector) or stick with just vector.
* Any pre-processing steps you found critical for medical or scientific texts.

I really want to make this useful for the team, but I'm stuck on making the retrieval precise enough. Any guidance from your experiences would be a lifesaver!

Thanks!



   
Quote
(@aiden22)
Reputable Member
Joined: 2 months ago
Posts: 350
 

Medical text is a worst case for naive chunking. Your problem is likely paragraph boundaries, not size.

For dense PDFs:
* Use sentence-transformers with medical domain models (like BioBERT) instead of generic embeddings.
* Chunk at 400 tokens max with 50 token overlap, but only split on hard paragraph breaks or section headers. Never cut mid-paragraph.
* Mandatory pre-processing: Extract and preserve figure/table captions as their own metadata-rich chunks. LLMs often rely on these.
* Always use hybrid search with a boosted keyword field. Medical terms have low semantic variance - "myocardial infarction" won't be found by "heart attack" embeddings alone.

The hybrid search in Relevance will save you. Tune the keyword weight heavily toward exact matches.


Show me the bill


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 2 months ago
Posts: 228
 

Great point about the paragraph boundaries, that tripped me up for weeks. The exact chunk size matters less than respecting the natural document structure.

Your hybrid search tip is spot on, but I'd add that you need to curate that keyword field. Auto-extraction from the PDFs gives you noisy junk. We built a simple allowed-terms list from MeSH headings first, then populated the field. It cut our false positives in half.

BioBERT is solid, but for truly niche sub-fields, we got better results fine-tuning a general model on a few hundred of our own documents. The initial embedding accuracy jump was worth the weekend of training time



   
ReplyQuote
(@infra_skeptic_9)
Prominent Member
Joined: 7 months ago
Posts: 602
 

Fine-tuning on a few hundred documents is interesting, but I'm stuck on the economics of it. You spent a weekend training, but did you track the compute cost? For 100k documents, that initial accuracy jump might be real, but you're now locked into maintaining that custom model pipeline forever. Every library update, every framework change becomes your problem.

And while a MeSH-based allowed-terms list is clever, it's another manually curated artifact. That list will rot. New drugs, new procedures, new acronyms appear constantly. Who's in charge of updating it when the research focus shifts next quarter? You've traded one type of noise for a maintenance time bomb.

It feels like we're all just building very elaborate, very expensive card catalogs.


Your k8s cluster is 40% idle.


   
ReplyQuote