Just migrated a client's internal Q&A pipeline from OpenAI to Perplexity API. Needed the web search context. Budget was supposed to be *better*.
Spoiler: It wasn't. Not even close.
Here's the math that made my eyes twitch. Pipeline does ~5000 queries/day, mostly `sonar-pro` (their "reasoning" model). Each query uses search, so it's 2x API calls: one for search, one for the actual completion. That's 10k API calls daily.
Perplexity's pricing page looks clean, but the reality hits different:
- `sonar-pro` is $5.00 per 1k *output* tokens.
- `sonar` (search) is $0.20 per 1k *input* tokens.
Our average query: 500 input tokens for search, 1500 output tokens for the answer.
```python
# Daily cost calc
search_calls = 5000
completion_calls = 5000
search_cost = search_calls * (500/1000) * 0.20 # $500/day
completion_cost = completion_calls * (1500/1000) * 5.00 # $37,500/day
total_daily = search_cost + completion_cost # $38,000/day
monthly = total_daily * 30 # ~$1.14M/month
```
Yeah. A million bucks a month. For 5000 queries.
OpenAI with GPT-4 + manual web search scraping was a fraction of this. The "per-output-token" model for the reasoning engine is a budget nuke for any real volume. Back to drawing board. Or maybe just building our own rag with web scrapes.