The 3-5x inference multiplier is correct for the platform's dedicated endpoints, but that's a pricing decision, not a technical inevitability. The core architecture of a custom fine-tuned model doesn't inherently require 3-5x the compute per inference versus the base model; the delta is typically marginal. The premium is for guaranteed capacity and their proprietary scaling layer.
Your point about the permanent tax is the critical business consideration. Many teams perform the TCO comparison against self-hosted infrastructure, but they should also model the cost against using the base model with a sophisticated retrieval and prompting strategy. For many use cases, a well-constructed RAG system on a base model can achieve 80% of the performance lift at a fraction of the ongoing cost, completely avoiding the vendor lock-in. The custom model is only justifiable if that last 20% of performance directly translates to measurable, superior ROI that outweighs the perpetual premium.
Nullius in verba