Skip to content
Notifications
Clear all

Switched from LlamaIndex to Weaviate for vector storage - better for hybrid search?

6 Posts
6 Users
0 Reactions
31 Views
(@annab)
Reputable Member
Joined: 3 months ago
Posts: 349
Topic starter   [#13267]

Hi everyone, I've been using LlamaIndex for about six months to manage document retrieval for our marketing content and customer FAQs. I liked how it integrated with our HubSpot data, but I kept hitting a wall with search relevance when users mixed keyword and semantic queries.

Last month, I decided to switch the vector storage backend from LlamaIndex's default to Weaviate. The main driver was needing more robust hybrid search—combining BM25 keyword scoring with vector similarity—which I heard is a core strength of Weaviate. I'm using it with the same OpenAI embeddings.

Initial results are promising. I'm seeing better recall for very specific, jargon-heavy marketing terms while still catching the conceptual questions. The setup was a bit more involved, especially around schema definition, but the queries feel more accurate.

Has anyone else made a similar switch for production use, especially in a marketing automation context? I'm curious about long-term maintenance, cost implications at scale, and if you've found any pitfalls in the integration layer. Also, I'm still figuring out the optimal weighting between keyword and vector search for our use case—any insights on tuning that would be appreciated.



   
Quote
(@charlesb)
Reputable Member
Joined: 3 months ago
Posts: 295
 

I run analytics infrastructure for a mid-market ecommerce platform, and we've had both LlamaIndex and Weaviate in production for different teams over the last two years.

**Hybrid Search Control**: Weaviate wins on tunability. You can adjust the weighting between BM25 and vector search per query. LlamaIndex's abstraction makes that harder. The sweet spot for our product docs landed at 0.6 vector, 0.4 keyword, but you'll need to A/B test.
**Cost at Scale**: Weaviate's managed cloud starts simple but watch the storage. Vector storage is roughly 2-3x the price of the object storage you'd use with LlamaIndex's typical setup. For us, that translated to ~$1200/mo for 500GB vs. ~$400/mo for Pinecone, which was the other candidate.
**Schema Rigidity**: This is Weaviate's main trade-off. You have to define your schema upfront, which took my team a solid week to model correctly. LlamaIndex is far more forgiving for rapid iteration. If your document structure changes often, that's a maintenance tax.
**Vendor Lock-in**: This is subtle. With LlamaIndex, you can swap vector stores. We moved from Pinecone to a self-hosted option with a few config changes. Weaviate is the database, so leaving means a full data migration. You're buying into their ecosystem.

I'd recommend Weaviate if your schema is stable and hybrid search precision is your primary KPI. If you're still experimenting heavily or need maximum flexibility to switch underlying components, LlamaIndex with a separate vector DB gives you an out. To be sure, tell us how often your document structure changes and if you have a hard monthly infra budget.


Beware of free tiers


   
ReplyQuote
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
 

Glad to hear the switch is working for you! I ran into a similar wall with hybrid search on a customer support FAQ project. That schema rigidity you mentioned gets easier, but wait until you need to update your vectorizer or add a new property mid-stream. 😅

For tuning, we found the optimal weighting swung wildly based on query intent. We ended up implementing a simple classifier upfront: if the query has specific product names or error codes, we skew heavily towards BM25; for "how do I" or conceptual questions, we go mostly vector. Weaviate's API makes that split pretty clean to implement.

On cost, watch your indexing operations, not just storage. If you're frequently syncing fresh HubSpot content, those re-indexes can add up on a managed Weaviate cloud plan. Might be fine for your volume, but it snuck up on us.


ship it


   
ReplyQuote
(@jakew)
Estimable Member
Joined: 3 months ago
Posts: 86
 

>We ended up implementing a simple classifier upfront

That's a fantastic point, and something we've been noodling on too. Our case is a bit of the opposite, funny enough - our marketing jargon is often conceptual (think "thought leadership" or "engagement loop") and our users' conceptual questions often contain those same terms. The classifier we're prototyping uses query length and the presence of stop words as a cheap heuristic first, then falls back to a tuned hybrid weight. It's messy, but it's cutting down on the really off-base results.

You're dead on about indexing operations costing more than you'd think. We batch our HubSpot syncs to a weekly job now after getting a surprise bill. The schema updates... yeah, that's the real pain. We had to add a new "content_region" property last week and the migration script felt like performing open-heart surgery.


Spreadsheets > opinions


   
ReplyQuote
(@consultant_mark)
Reputable Member
Joined: 5 months ago
Posts: 231
 

Your point about a classifier for query intent is spot on, and it aligns with a principle we enforce in revenue operations: the scoring logic must match the business process. For a sales team searching product documentation, the optimal weighting for a query from an AE in a discovery call is different from one from a support agent handling a ticket.

We found that static tuning, even with a good heuristic, eventually decays as product language evolves. We now bake the classifier output, the specific search context, and the result's click-through rate back into our tuning parameters every quarter. It's a lightweight feedback loop that prevents the weights from drifting.

The real hidden cost you mentioned, indexing operations, is often a total cost of ownership blind spot. Beyond batching, we had to implement a versioning flag in our schema to perform a zero downtime migration when adding a property, because a full re-index of our historical CRM data would have taken 18 hours and stalled the sales team.



   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

Glad the initial results look good, but you're in the honeymoon phase. "More robust hybrid search" is the marketing line. The real question is, how does it handle when your marketing team invents a new buzzword next quarter and your schema can't handle it without a migration?

The promise of tunable weighting is a siren song. You'll spend more time A/B testing and building query classifiers than you ever did wrestling with LlamaIndex's abstractions. And wait until you get that first bill for re-indexing your HubSpot content after a major campaign. The storage cost is one thing, the operational cost of keeping it all in sync is the real tax.

Schema definition being "a bit more involved" is the understatement. That rigidity will bite you the first time compliance asks for a new audit trail property on your stored vectors.


Trust but verify


   
ReplyQuote