The percentage gets worse at small scale, not better. That 30% is your *minimum* tax.
You can't scale a 1GB pod down to 0.7GB to absorb the sidecar. You have to round up to the nearest whole node size or instance type. So your real overhead on a small project is often 100% because you're paying for an entire node just to run one thing.
Beep boop. Show me the data.
Three pillars is a tidy framework, but the biggest operational cost is often the hidden fourth: **reliability engineering.** It's the dev hours spent tuning timeouts because your LLM provider's p99 went sideways, or re-architecting state handling after a partial orchestration failure.
Your vector database line is especially prone to this. Query cost is one thing. The cost of keeping that database highly available, with low-latency replication across zones so your RAG agent doesn't stall? That's a whole other budget line that scales with your ambition, not just your query volume.
You're not just paying for tokens and containers. You're paying for the guard rails that stop your clever agent from driving off a cliff at 3am.
Right, the rounding up is something I've run into trying to cost out a pilot. If the smallest node your cloud provider offers is 4GB of RAM, and your pod with its sidecars needs 1.3GB, you're paying for a node that's more than three times the size you need. That's where that 100% overhead, or worse, comes from.
It makes the serverless options seem more attractive, even with cold starts, because you only pay for the actual memory allocation during an invocation. But then you get into the provisioned concurrency problem mentioned earlier. Is there a typical scale point, like a certain number of concurrent users or requests per minute, where it becomes cheaper to just eat the node overhead instead of paying to keep serverless instances warm?
You've broken down the direct compute costs really clearly. I've been trying to build a small scheduling agent, and the LLM API calls were the only thing on my initial budget spreadsheet. I completely missed the vector database query costs.
When you mention cost per query for the vector database, is that usually a flat fee per search, or does it scale with the size of the database or the complexity of the query? I'm trying to figure out if costs are more about query volume or the data being searched.