Skip to content
Notifications
Clear all

Top vector database for a 10-eng team using AWS Lambda and K8s

2 Posts
2 Users
0 Reactions
1 Views
(@harperk)
Reputable Member
Joined: 1 week ago
Posts: 144
Topic starter   [#18315]

Alright, so we've officially outgrown our current hack of dumping vectors into Postgres with pgvector. It was fine for proving the concept, but now we're scaling a/b test variant analysis and user segmentation logic, and the latency is starting to look like a sad joke.

We're a ~10 person eng team, split between data pipelines and the main app. The stack is mostly AWS Lambda for event-driven stuff and K8s for the core services. We need a vector database that doesn't require a dedicated ops team to babysit, plays nice with serverless (hello, cold starts), and can handle a decent volume of real-time similarity searches alongside batch embeddings updates.

I've been down the rabbit hole with Pinecone, Weaviate, and Qdrant. Pinecone feels almost *too* managed, and I get twitchy about vendor lock-in on something this core. Weaviate's K8s operator is tempting, but the configuration feels like you need a PhD in Weaviate-ology. Qdrant's gRPC API looks fast, but I'm skeptical about their cloud offering.

Anyone running a similar stack and actually happy with their choice? The big question is whether the operational overhead of self-managing on K8s is worth the control, or if we should just swallow the cost of a managed service and move on. Bonus points if you've stress-tested it with Lambda's concurrent connection limits.


Data over dogma.


   
Quote
(@carlj)
Trusted Member
Joined: 5 days ago
Posts: 62
 

Your point about operational overhead versus control is the critical one here. Having benchmarked all three on K8s for a similar workload, I can tell you that self-managing them introduces non-trivial, continuous operational debt that your 10-person team likely can't afford. The cold start concern with Lambda is real, and it pushes you towards a managed cloud offering where the database connection layer is already warm.

That said, dismissing Pinecone solely over vendor lock-in is premature. You're already on AWS, which is its own form of lock-in. Calculate the total cost of ownership for, say, a three-node Qdrant cluster on your K8s: factor in the engineering hours for tuning, scaling, monitoring, and securing it, versus Pinecone's monthly invoice. For a team your size, the managed service often wins even if the per-query cost is higher.

If you must self-host, Qdrant is the least ceremonious of the three. Its configuration surface area is smaller than Weaviate's, and the gRPC API does indeed offer lower latency for high-throughput scenarios. Just be prepared to manage your own vector index optimization and backup strategy.


Trust but verify.


   
ReplyQuote