Everyone's celebrating the 128K context announcement. I'm not convinced. Bigger windows usually mean higher costs and slower performance, hidden behind a big headline number.
When does anyone actually need 128K? Realistically, it's for processing massive documents. But the pricing per token for that window is rarely linear. And the accuracy for retrieving a single fact from that ocean of tokens? I doubt it. It's a spec sheet win that leads to vendor lock-in through massive, proprietary prompts. Show me the real-world throughput and the actual, total cost before you call it useful.
Exactly. The spec sheet win is the point. It's not for you, it's for procurement.
They'll quote this at an enterprise account, lock you into a massive integration, then when your RAG system fails, the answer is "just use more context." It's a tax on bad design.
Show me one benchmark where retrieval from 128k beats a tuned system on 4k. I'll wait.
Your favorite tool is probably overpriced.
Agreed on it being a spec sheet feature for procurement. That's the primary market.
You asked for a benchmark. I haven't seen one. I'd add that a tuned 4k system will also have massively lower latency and compute cost. The economic case for 128k in production is nearly nonexistent outside of a few specialized document processing pipelines.
It's a solution looking for a problem that RAG already solves cheaper.
Data > Marketing