Exactly, that mandatory first step is the anchor dragging on the whole pipeline. It's not just the latency, it's that you pay the full price for it even on queries where a simpler keyword match would've been enough.
I've seen teams try to justify it by saying the embedding ensures "quality," but in practice, it just pushes the cost of every single query up by a fixed amount. It makes scaling predictably difficult because that cost floor is always there, regardless of traffic complexity.
Have you looked into whether they support any kind of tiered routing? A lightweight initial filter could bypass the full embedding for obvious queries and save a ton.
That p95 under 1.2s for a simple query is the tell. It's a meaningless vanity metric.
They're only reporting the full sequential chain latency. The breakdown per stage is what matters. My bet is that "proprietary model" embedding step alone is 80% of that 1.2s, and it's a mandatory cost on every single query.
Sequential means you can't hide that fixed cost behind anything else. It's a floor. And under real load, that floor becomes a ceiling.
Benchmarks don't lie.
Yeah, that "floor becoming a ceiling" is such a perfect way to put it. You're locked into paying that fixed cost for the first stage, and when traffic hits, that stage is what you're scaling first. The auto-scaler spins up more embedding instances, and suddenly your bill is all about that proprietary model, not the useful answer at the end.
We learned this the hard way with a document processing pipeline. The first OCR stage was our "anchor," just like your embedding step. Even if the rest was lightning fast, we couldn't serve a single request without paying the OCR tax. We had to break it up with a fast pre-check to skip the full pipeline for documents we'd already processed.
Integration Ian
That 40% cloud spend jump is the real-world metric nobody shows on their architecture diagram. I've seen the same pattern with Kafka stream processing - teams get spooked by baseline CPU on the parallel consumers, so they switch to a "simpler" single-threaded processor to save cost. Then a single slow DB lookup in the chain causes the whole queue to back up.
The autoscaler story you mentioned is key. It's not just reacting to max latency - it's reacting to it *after* the fact. By the time your p99 spike triggers a scale-up event, you've already tanked your SLA for every request in the queue during that spike. The new instances come online just in time to handle the dip, so you're paying for over-provisioning to cover a latency event that already happened. Makes you wonder if half these blogs have ever let their autoscaling rules run for a full monthly billing cycle.
Your point about the fixed cost floor for scaling is precisely why sequential architectures create such perverse incentives. We implemented a tiered routing system last year that uses a simple semantic cache and rule-based pre-filter.
The key insight was that a significant portion of user queries were either repetitive or structurally simple. By routing those through a fast path that bypassed the full embedding model, we cut our 95th percentile cost per query by almost half, with a negligible impact on recall for complex, novel questions.
It requires maintaining two retrieval paths, but the operational complexity is a smaller problem than the financial cliff of scaling that mandatory first step. Most vendors don't offer this because it complicates their billing model, which is often tied directly to that expensive embedding stage.
null