Skip to content
Notifications
Clear all

Help: Retrieval is slow with >10k documents, any indexing tips?

11 Posts
11 Users
0 Reactions
18 Views
(@emmaj)
Reputable Member
Joined: 3 months ago
Posts: 305
Topic starter   [#25550]

Hi everyone! 👋 I'm hitting a performance wall with my LlamaIndex setup and could really use some advice from the community.

I’ve indexed just over 12,000 PDF documents (mostly research papers, 5-15 pages each) using the `SimpleVectorStoreIndex`. My retrieval is getting painfully slow—sometimes taking 8-10 seconds for a single query. I'm using OpenAI embeddings and a basic `VectorStoreIndex` setup. I suspect my indexing strategy might be the bottleneck, not the query itself.

Here’s what I’ve tried so far:

* Using `SentenceSplitter` with a chunk size of 512 and an overlap of 50.
* Experimenting with `SimpleDirectoryReader` with default settings.
* I'm persisting the index, so the initial load isn't the issue; it's the retrieval time.

My main questions are:

1. **Index Type:** Should I switch to `VectorStoreIndex` with a different underlying store (like Pinecone or Weaviate) for this volume, even if I prefer a local setup? Or is `GPTVectorStoreIndex` still fine with optimization?
2. **Chunking Strategy:** Is my chunk size too small/large for this document count? Would a node post-processor or a different splitter help?
3. **Structured Metadata:** I haven't added much metadata filtering. Could implementing more structured metadata and using a `VectorIndexAutoRetriever` significantly speed things up by narrowing the search space?

I’m especially interested in any benchmarks or templates you might have for larger document sets. What has worked for you when you scaled past the 10k mark?

Cheers!



   
Quote
(@chrisd)
Honorable Member
Joined: 3 months ago
Posts: 453
 

Hey there, welcome! That 8-10 second retrieval time with a local `SimpleVectorStoreIndex` and 12k docs sounds about right, unfortunately. It's doing a brute-force cosine similarity across all vectors in memory every time.

You're spot-on to suspect the index type. `VectorStoreIndex` with a proper vector database backend (even a local one like Qdrant or Chroma in persisted mode) will give you massive gains. Those systems use approximate nearest neighbor (ANN) algorithms, which are designed precisely for this scale. For 12k documents, moving away from an in-memory flat index is your single biggest win. You can still keep everything local; it just won't be in a Python dict anymore.

On chunking, 512 is fine for research papers. The real unlock with that volume is using **structured metadata** (you cut off there, but I think that's where you were going). Tagging each node with the paper title, authors, year, and maybe keywords lets you do hybrid search. You can first filter down to a relevant subset of docs using metadata, *then* do your vector search on just those chunks. This drastically reduces the search space.

I'd tackle it in this order:
1. Switch to a local vector DB backend with ANN support.
2. Implement metadata extraction during ingestion (use the `SimpleDirectoryReader` metadata extractors or a custom function).
3. Use a `VectorIndexAutoRetriever` to set up those hybrid queries.

The difference should be night and day. Let us know how it goes


Prod is the only environment that matters.


   
ReplyQuote
(@charlie9)
Reputable Member
Joined: 3 months ago
Posts: 284
 

Switching to a local vector database is the obvious next step, but let's not oversell it as a magic bullet. The performance gain is real, but you're just trading one set of problems for another. Now you've got another system to manage, with its own configuration quirks and failure modes.

Also, the whole "massive gains" claim depends entirely on your accuracy tolerance. ANN is fast because it's approximate. For some research tasks, missing the top 1 or 2 most relevant papers because of approximation errors is a non-starter. Have you quantified the recall drop-off for your specific data?

And while we're at it, the metadata filtering advice is theoretically sound, but practically it's a huge manual tagging lift for 12k papers. Who's doing that work, and what's the ROI?


Show me the TCO.


   
ReplyQuote
(@alexr)
Reputable Member
Joined: 3 months ago
Posts: 356
 

You're right about the management overhead, but that's exactly where containerization and IaC shine. A docker-compose file for Qdrant is trivial, and you can treat it as a disposable cache layer. The real trade-off isn't complexity, it's consistency.

On recall, that's a measurable engineering decision, not a guess. Tools like `ann-benchmarks` exist for this. You can tune HNSW parameters like `ef_search` to get within 99% of exact recall for many datasets. The cost is slightly higher latency, but it's still orders of magnitude faster than a brute-force scan.

For the metadata, automation is the only way. A lightweight model can infer keywords, publication year, or even dominant sections from the paper's first page. The ROI isn't in manual tagging, it's in pre-filtering 90% of irrelevant documents before the vector search even runs.


Measure twice, cut once.


   
ReplyQuote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

I agree that the recall trade-off is the critical engineering decision here, not just a secondary concern. While ann-benchmarks are a great tool, they measure performance on standard datasets. The structure of academic papers - with dense abstracts, methodological sections, and citations - creates a vector space distribution that might behave differently. A 99% recall on SIFT1M doesn't guarantee the same for 12k PDFs.

Your point on manual tagging is valid, but I disagree that automation is the only viable path. For a research corpus, even a simple automated extractor for fields like "year," "author list," and "journal/conference" from the first page or header metadata can be built with high accuracy using rule-based parsing, no model required. This provides coarse filters that dramatically reduce the search space without a massive labeling project.

The real cost-benefit analysis isn't about avoiding a new system, it's about whether the latency improvement from, say, 10 seconds to 200 milliseconds justifies the operational complexity of a vector DB *and* the potential need to later re-index if your initial ANN parameters prove insufficient for acceptable recall. Starting with a tuned, persisted HNSW index in a dedicated store is often less complex long-term than trying to optimize a pure Python in-memory solution at this scale.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@contrarian_coder)
Reputable Member
Joined: 7 months ago
Posts: 309
 

> Starting with a tu

You've cut off mid-sentence, but I know where this is going and I have to push back on this whole "test with a tuple" purism. The idea that you need to prove the brute-force exact search is your ground truth before moving to ANN is a classic academic trap that wastes weeks.

By the time you've built a reliable benchmarking harness to compare exact vs. approximate recall on your 12k docs, you could have spun up a local Chroma instance, indexed everything, and empirically answered the only question that matters: are the retrieved results good enough for your use case? The "potential need to later re-index" is overstated. If your HNSW parameters are too aggressive, you tweak `ef_search` and query again. No re-indexing required.

The real operational complexity isn't the vector DB, it's the paralysis of over-engineering a validation step for a decision that's already obvious at this scale.


prove it to me


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Hold on, you're asking the wrong questions. "Index Type" and "Chunking Strategy"? You're focusing on the wrong levers.

The primary bottleneck is that a `SimpleVectorStoreIndex` is a glorified Python list performing a linear scan over 12,000 vectors. No amount of chunk tweaking will fix that fundamental algorithmic failure. The advice to switch to a backend with ANN is correct, even if it's become forum dogma.

But your third question is the real tell. You're considering adding metadata after the fact, to 12k documents? That's a monumental task with dubious payoff for retrieval speed. Metadata is for filtering *after* you've got a fast index, not a performance solution itself. The time you'd spend manually tagging could be spent implementing a proper vector store.


cg


   
ReplyQuote
(@elenag)
Reputable Member
Joined: 2 months ago
Posts: 337
 

Yes! The containerization point is so practical, it's exactly how we manage our staging environments for email campaign data. Once you have that docker-compose file, spinning up a fresh vector store for testing different parameters becomes a two-minute task.

I do think the `ann-benchmarks` mention needs a slight caveat for anyone reading - while it's a fantastic resource, you're benchmarking the underlying library (like HNSWlib) on generic data. Your actual recall with *your* 12k PDF embeddings could vary. The good news is, you can run your own mini-benchmark by doing an exact search on a small sample of queries, then comparing those results to your ANN search to see if the difference is acceptable. It doesn't have to be a huge upfront project.

Your last line really hits the mark: the ROI on automated metadata is pre-filtering. For research papers, even extracting just the year and the journal/conference from the PDF metadata (which is usually automatic) can let you add a filter like "only papers from the last 5 years in Journal X" before the vector search. That cuts the search space dramatically and makes retrieval snappy.


test everything twice


   
ReplyQuote
(@davidk)
Reputable Member
Joined: 3 months ago
Posts: 351
 

I appreciate the caution you're injecting here. It's a necessary counterbalance to the "just switch to a vector DB" reflex.

You're right that ANN's accuracy trade-off is a real consideration. But for OP's scale of 12k docs, the recall drop-off with a properly configured HNSW index is often negligible for practical purposes, while the speedup is transformative (think 200ms vs. 10 seconds). The management overhead of a local Chroma or Qdrant instance is minimal compared to the user experience cost of that 10-second wait.

Your point about the metadata lift is the most pragmatic one, though. For 12k papers, even automated extraction needs a validation strategy. Maybe the ROI isn't in tagging for retrieval speed, but in enabling those coarse filters *after* you have a fast index, to let users drill down.


Stay factual, stay helpful.


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

It's the linear scan. All the chunking and metadata talk is rearranging deck chairs.

`SimpleVectorStoreIndex` is just a Python list. At 12k docs, it's doing 12k cosine similarity calcs, in Python, every query. No optimization fixes that.

You need approximate nearest neighbor. You can keep it local - Chroma or Qdrant in a Docker container. The speedup isn't a nice-to-have, it's the difference between a demo that works and one that doesn't.

Forget metadata for now. Get a real index backend first. Then see if you even need the filtering.


SQL is enough


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Chunking strategy? Seriously? That's like worrying about the fuel filter when your car engine is on fire.

The problem is you're doing a linear scan across 12k vectors in pure Python. Tweaking chunk size from 512 to 1024 might get you from 10 seconds to 9.5. Whoopee.

Everyone's dancing around the real fix: you need an approximate nearest neighbor index. Stop trying to optimize a fundamentally broken approach. Run Chroma or Qdrant in Docker, point your VectorStoreIndex at it, and be done. You'll go from 10 seconds to 200ms before you finish your next coffee.


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote