Skip to content
Notifications
Clear all

Help: Retrieval is slow with >10k documents, any indexing tips?

2 Posts
2 Users
0 Reactions
31 Views
(@liamj)
Trusted Member
Joined: 3 months ago
Posts: 34
Topic starter   [#4272]

I've been conducting a performance benchmark for a potential enterprise implementation of LlamaIndex and have hit a significant bottleneck. My current test corpus consists of approximately 15,000 PDF documents (primarily technical manuals and internal process documents, averaging ~10 pages each). The ingestion and indexing phase, while lengthy, is acceptable as a one-time cost. However, the retrieval latency during query time is consistently above 4-5 seconds for a single query, which is untenable for a responsive application. This is using the standard `VectorStoreIndex` with OpenAI's `text-embedding-ada-002` and a basic `SimpleNodeParser` with a chunk size of 1024.

I have already ruled out the embedding API call as the primary culprit through isolated testing. The slowdown appears to be intrinsic to the similarity search operation over the vector store. I am using the default in-memory `SimpleVectorStore` for prototyping. My hypothesis is that the sequential search over a flat index of over ~150,000 text chunks (nodes) is the root cause.

Before I proceed with a significant infrastructure shift, I want to exhaustively explore configuration and architectural options within the LlamaIndex framework itself. My specific questions for the community are:

* **Indexing Strategy:** Has anyone performed a comparative analysis of different index types (`VectorStoreIndex`, `TreeIndex`, `KeywordTableIndex`) at this scale? The literature suggests a `VectorStoreIndex` with a more sophisticated backing store is the path, but I require empirical data or case studies.
* **Node Parsing & Chunking:** Could the chunk overlap or chunk size be a contributing factor to search inefficiency, perhaps by creating an overwhelmingly large and redundant index? What is the recommended heuristic for these parameters with large, dense documents?
* **Vector Store Selection:** This seems the most likely avenue for improvement. For those in production:
* What is the observed performance delta between the in-memory store and integrated options like `PineconeVectorStore`, `WeaviateVectorStore`, or `QdrantVectorStore` at a comparable scale?
* Are there any non-obvious trade-offs in accuracy (recall/precision) when moving to these managed services?
* Is there a compelling argument for a local, persistent store like `Chroma` or a local `Faiss` index for latency control, assuming the hardware is provisioned adequately?
* **Metadata Filtering:** I have implemented some basic metadata (document source, date). Are advanced filtering techniques or the use of `IndexStruct` types practically useful for pre-filtering the candidate pool before the full vector search, thereby reducing the effective search space?
* **Systematic Evaluation:** Beyond anecdotal "it feels faster," what methodology have you used to quantitatively measure retrieval performance (latency p95/p99, throughput) before and after optimizations?

My goal is to establish a reproducible, optimized baseline. I am less interested in "try this service" suggestions without an accompanying analysis of the performance characteristics and the underlying mechanism for the improvement. Any detailed workflow reports, configuration snippets (focusing on parameters like `chunk_size`, `similarity_top_k`, `vector_store_query_mode`), or links to rigorous benchmarking would be immensely valuable.

—LJ


—LJ


   
Quote
(@mattl88)
Active Member
Joined: 3 months ago
Posts: 8
 

Good catch on isolating the embedding API. Your hypothesis about the sequential scan in `SimpleVectorStore` is correct. It's a brute-force cosine similarity calculation over all vectors, which becomes a real bottleneck once you cross a few thousand nodes.

Before switching infrastructure, you should first confirm the bottleneck is the similarity search itself, not something else in your query pipeline. Add some simple timing around the `_vector_store.similarity_search` call. If that's consistently taking 3+ seconds, then yes, you need a proper vector index.

The most direct next step within LlamaIndex is to test with their `GPTVectorStoreIndex` using a different underlying store that supports approximate nearest neighbors (ANN). Try using the `PineconeVectorStore` or `WeaviateVectorStore` integrations, even just for a benchmark. The latency difference between exact and approximate search at your scale should be dramatic, likely bringing you under 500ms.


Measure twice, migrate once.


   
ReplyQuote