Skip to content
Notifications
Clear all

Has anyone benchmarked inference speed between PlayHT and WellSaid Labs?

10 Posts
10 Users
0 Reactions
24 Views
(@hiroshim)
Noble Member
Joined: 3 months ago
Posts: 767
Topic starter   [#24120]

I am currently evaluating text-to-speech services for a latency-sensitive application involving dynamic content generation, and the inference speed—specifically the time from API call to receipt of the first audio byte (time-to-first-byte, or TTFB) and the total time for full audio synthesis—is a critical performance metric. While both PlayHT and WellSaid Labs produce high-quality output, their architectural approaches (as inferred from documentation) suggest potential differences in pipeline efficiency that would materially impact user experience in interactive scenarios.

I have conducted preliminary, informal tests using their respective standard APIs, but I require more rigorous, comparable data. My initial methodology involved:
* **Environment:** A dedicated cloud instance (us-east-1) to minimize network variance.
* **Sample Text:** A standardized 250-character paragraph of neutral prose.
* **Measurement:** Scripts using `curl` with the `-w` flag for timing details, and custom Python clients using `time.perf_counter()` to capture client-side latency.
* **Parameters:** Default voice, 24kHz mono output format for both services.

My initial observations, which should not be considered definitive, indicated a variance. For example:

```bash
# Simplified measurement approach for TTFB
curl -X POST "https://api.play.ht/v2/tts"
-H "Authorization: Bearer $API_KEY"
-H "Content-Type: application/json"
-d '{"text":"$SAMPLE_TEXT", "voice":"$VOICE_ID"}'
-o output.mp3
-w "time_namelookup: %{time_namelookup}ntime_connect: %{time_connect}ntime_appconnect: %{time_appconnect}ntime_pretransfer: %{time_pretransfer}ntime_starttransfer: %{time_starttransfer}ntime_total: %{time_total}n"
```

The `time_starttransfer` value here approximates TTFB. However, this single-point measurement is insufficient.

I am seeking community input to validate or expand upon these findings. A comprehensive benchmark would need to control for and report on:
* **Regional API endpoints** and their effect on latency.
* **Voice-specific performance,** as some voices may utilize different underlying models.
* **Concurrent request handling** and any scaling latency penalties or efficiencies.
* **The impact of text length** on the linearity of synthesis time.
* **Streaming response performance** versus monolithic file delivery.

Specifically, I am interested in whether anyone has performed systematic A/B testing under controlled conditions, accounting for:
1. Cold vs. warm model inference latency.
2. Network hop consistency (using tools like `mtr`).
3. The performance tier of the subscribed plan (e.g., pay-as-you-go vs. enterprise).

Subjective feedback on "speed" is less useful without the accompanying methodology. If you have undertaken such measurements, sharing your approach, sample size, and results—even in raw form—would be invaluable. Furthermore, insights into the internal architecture of either service (e.g., pre-loading of voice models, caching strategies, or use of speculative execution) that you may have gleaned from support documentation or engineering blogs would help explain observed performance characteristics.



   
Quote
(@connork)
Reputable Member
Joined: 2 months ago
Posts: 216
 

That's a really interesting approach. When you say "initial observations," were you seeing a consistent winner between them on TTFB, or was it too variable? I'm curious because I've only done much simpler timing checks for my own use. Your setup sounds way more controlled.



   
ReplyQuote
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Too variable to call with any confidence. Their latency isn't a constant, it's a distribution you're fighting with. The "winner" on any given run depends more on which service's queue you hit at a bad time.

For dynamic content, you're better off benchmarking the 95th or 99th percentile, not the average. That's where the user experience dies.

My informal p95s were close enough that I wouldn't pick one over the other on speed alone. Quality differences and cost were bigger factors.


Prove it.


   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Good point on p95/p99 over average. But if their distributions are that similar, the performance is effectively a coin flip.

That makes the decision a lot easier. Pick the cheaper one. Or the one whose voice you can stand listening to for the next year.

Complex benchmarks for a negligible difference is overthinking it.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@evanj)
Estimable Member
Joined: 3 months ago
Posts: 189
 

I'm still in my own evaluation phase, and I definitely get what you mean about simpler timing checks. My experience mirrors that variability.

Even in my controlled tests, the results aren't stable enough to crown a "winner." One service might be 200ms faster on three consecutive calls, then suddenly be a full second slower on the next. It feels like there's a hidden queue or load factor that the public metrics don't capture.

What does your simpler timing setup look like? I'm wondering if I've over-engineered my own and should strip it back to just measuring the core request.



   
ReplyQuote
(@backend_latency_queen)
Honorable Member
Joined: 4 months ago
Posts: 613
 

That hidden queue feeling is exactly why I advocate for measuring p95/p99, not just average latency. Averages smooth over those punishing, inconsistent delays.

Your timing setup is likely fine, but you need to run it hundreds of times to see the shape of the distribution. A single-second delay could be the 99th percentile event you're trying to avoid. I'd suggest a small Go script that logs each request's duration, then runs a quick analysis on the results file. That's simpler than trying to build something that reacts to each individual call.

Strip it back to the core request, but run it as a batch job overnight. The volume of data is what reveals the real performance characteristic.


sub-100ms or bust


   
ReplyQuote
(@ellawest)
Estimable Member
Joined: 2 months ago
Posts: 102
 

Hundreds of calls is a great starting point, but it's still just a snapshot. I'd push back on treating that as a definitive profile. Unless you're running those hundreds of calls at the same time every day, across different days of the week, you're only capturing the service's behavior under your specific, narrow conditions.

A batch job overnight gives you one pattern. What happens during your region's business hours? What happens when the vendor pushes a model update, or has a regional hiccup? The p95 you measure tonight could be meaningfully different from the p95 next Tuesday at 2 PM.

Rigorous benchmarking for a production decision means accepting you're measuring a moving target. Your script should be a permanent fixture, not a one-off test, or you're just getting a comfortingly precise number for a transient state.


audit logs don't lie


   
ReplyQuote
(@charlotte0)
Reputable Member
Joined: 3 months ago
Posts: 241
 

Your methodology is a strong starting point for isolating inference speed, especially using a dedicated instance and neutral sample text. I've been setting up similar tests.

To build on that, have you considered controlling for voice selection bias? Some voices, even default ones, might use different underlying model architectures that could affect latency. You might want to test across a shortlist of 2-3 common voices per service using your same script.

Also, for dynamic content, how are you factoring in the length variability of real-world inputs? A 250-character benchmark is good, but the relationship between character count and synthesis time isn't always linear.



   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

I agree with the spirit of this, but calling it a coin flip is a bit reductive. If the p95s are truly neck and neck, it absolutely shifts the focus to other factors like cost and voice preference.

However, those "negligible" differences at the 95th percentile can still represent hundreds of milliseconds of jitter that are perceptible in a truly interactive app. The decision point is whether that jitter falls within your specific application's tolerance window. For some use cases, it *is* negligible. For others, it's the difference between feeling instant and feeling laggy.

So it's less about overthinking complex benchmarks and more about knowing what your own threshold for "good enough" actually is. If both services clear that bar on speed, then your decision framework is spot on.


Keep it constructive.


   
ReplyQuote
(@ashp99)
Honorable Member
Joined: 2 months ago
Posts: 377
 

Exactly. A batch job is the simplest way to get that volume without over-engineering.

The trick is what you *do* with those hundreds of logs. A quick Python script with numpy or Pandas can spit out percentiles in seconds. No need to get fancy, just the core stats.

My only caveat: running it overnight might skew your numbers if the vendor does maintenance or has lighter load then. You could get a misleadingly good p95. Maybe schedule a few short batches across a 24-hour period?


data over opinions


   
ReplyQuote