I've been conducting a systematic evaluation of retrieval-augmented generation (RAG) pipelines built with LlamaIndex, specifically measuring the impact of embedding model choice on answer accuracy. During my controlled benchmarking, I've isolated a critical failure mode: **embeddings mismatch**, which leads to the retrieval of semantically irrelevant context and consequently generates "weird" or hallucinated answers. This isn't merely a qualitative observation; my latency and precision-at-k metrics clearly degrade when components are incongruent.
The core issue arises from an often-overlooked assumption: that the embedding model used for indexing documents operates in the same vector space as the one used—implicitly or explicitly—during query-time retrieval within the LlamaIndex stack. My typical test harness involves:
```python
from llama_index.core import VectorStoreIndex, SimpleDirectoryReader
from llama_index.embeddings.huggingface import HuggingFaceEmbedding
# Indexing with Model A
embed_model_a = HuggingFaceEmbedding(model_name="BAAI/bge-small-en-v1.5")
documents = SimpleDirectoryReader("./data").load_data()
index = VectorStoreIndex.from_documents(documents, embed_model=embed_model_a)
index.storage_context.persist(persist_dir="./index_a")
# Querying with a potentially mismatched setup
from llama_index.core import StorageContext, load_index_from_storage
# Scenario 1: Direct load, same embed model (control) - OK
storage_context = StorageContext.from_defaults(persist_dir="./index_a")
index_loaded = load_index_from_storage(storage_context, embed_model=embed_model_a)
# Scenario 2: Load persisted index but inadvertently use Model B - PROBLEM
embed_model_b = HuggingFaceEmbedding(model_name="sentence-transformers/all-MiniLM-L6-v2")
index_mismatched = load_index_from_storage(storage_context, embed_model=embed_model_b)
query_engine = index_mismatched.as_query_engine()
# This query will run, but retrieval is now flawed.
```
The symptoms are not always obvious. The system doesn't crash; it simply returns answers derived from incorrect source nodes, as the query embedding from Model B cannot correctly compute similarity against vectors generated by Model A. My benchmarks show a drop in Hit Rate from ~0.92 to near-random (~0.15) under such a mismatch.
I am seeking to compile a definitive list of all potential points of mismatch within LlamaIndex's architecture. My preliminary list includes:
* Persisting an index with one `embed_model` and loading it without explicitly specifying the same model.
* Using a different embedding model for a separate `ServiceContext` (or modern `Settings`) at query time versus index creation.
* Utilizing a vector store (e.g., Pinecone, Weaviate) that performs its own server-side embedding if the client-side `embed_model` is not correctly disabled or aligned.
* The more subtle case of using the same *model name* but with different normalization settings (e.g., one instance with `normalize=True`, another without).
Has anyone else performed rigorous ablation studies on this? I am particularly interested in reproducible, minimal code examples that demonstrate the failure and the correct patterns to enforce embedding consistency. Furthermore, what are the best practices for versioning embeddings alongside index artifacts to prevent silent degradation? The documentation mentions persistence, but the criticality of the embedding model as a primary dependency feels under-emphasized.
numbers don't lie
numbers don't lie
So you've got two different models writing vectors. That's like having separate address books for the same people. Of course your retrieval breaks.
But why do you need to swap models at query time at all? Are you trying to upgrade the index without re-indexing? That's a trap. Just re-index with the new model and be done with it.
All this complex benchmarking for a problem you can fix with a `docker run` and some patience.
Keep it simple
You're absolutely right about the vector space mismatch. It's a classic configuration drift problem, not unlike a security group referencing a deleted VPC.
Your test harness shows the exact setup. The critical line is when you initialize the `VectorStoreIndex`. If the query engine later uses a different model, even a newer version of the same name, you'll get silent failures.
One related caveat I've seen in production: teams sometimes use the same embedding model name but forget to pin the version across their indexing service and query service. A container update on one side can introduce the same mismatch. Consistent artifact versioning is key, just like locking down IAM policy versions.
security by default