We've been evaluating Perplexity's API for internal knowledge search across our engineering docs and incident reports. The promise is solid: real-time web context plus our own data. But in load tests, we're seeing inconsistent p95 latency spikes that make me question its readiness for high-volume enterprise use.
Our test setup:
- Simulated 50 concurrent users querying a mix of simple (keyword) and complex (natural language) questions.
- Ingested ~10GB of our internal Markdown docs via their data upload.
- Measured from our primary AWS region (us-east-1) to their endpoints.
The results weren't terrible, but they're not "enterprise search" stable either. Simple searches returned in ~800ms p50, but p95 ballooned to 4-5 seconds. Complex queries with web context occasionally hit 10+ seconds. For a team expecting sub-2-second responses, that's a non-starter.
Key bottlenecks we observed:
* **Web context retrieval** seems to be the wildcard. Enabling it adds unpredictable overhead.
* **Data source blending** (our docs + web) appears to be serialized, not parallelized, in many cases.
* Their rate limiting is aggressive; we hit throttling at volumes that wouldn't faze Elasticsearch or even OpenAI's APIs.
We're considering a hybrid approach: using Perplexity for complex, web-augmented queries only, and a local vector store for everything else. But that defeats the purpose of a unified API.
Has anyone else stress-tested this in a production-like environment? Specifically:
* What's your actual latency at >100 RPM?
* Did you find tuning parameters (timeouts, context limits) that helped?
* Is the latency cost from data upload searches predictable?
I'll share our full benchmarking Terraform config and k6 load test script if anyone wants to compare notes. Right now, my verdict is "not yet" for scaling.
-shift
shift left or go home