Skip to content
Notifications
Clear all

First-time evaluator - what metrics should I track in a trial?

12 Posts
12 Users
0 Reactions
21 Views
(@chloek4)
Reputable Member
Joined: 2 months ago
Posts: 303
Topic starter   [#24352]

Hey everyone! 👋 I'm deep into evaluating Synthesia for a potential integration into our content workflow, and I'm about to start a trial. I know from past experience with other APIs that it's easy to get dazzled by the demo and miss the practical, day-to-day metrics that really matter for reliability.

I want to set up a proper testing framework from day one. Beyond the obvious "video quality," what specific technical and operational metrics should I be logging?

My initial list includes:
* **API Response Times:** For both generation status checks and the final asset delivery webhook.
* **Webhook Reliability:** Percentage of successful payload deliveries vs. missed calls. I'll be testing with a dedicated endpoint that logs everything.
* **Error Rate & Types:** Categorizing `4xx` vs. `5xx` responses, and specific codes for quota issues or malformed requests.
* **Render Time Consistency:** Does a 1-minute video take roughly the same time at 2 PM as it does at 2 AM?
* **Cost per Unit:** Tracking credit/video length ratios against the stated pricing model.

I'm also curious about the **connector quality** for platforms like Zapier or Make. Has anyone stress-tested them? For example, I'd set up a scenario like this in Make to test error handling:

```json
{
"scenario": "Webhook Retry on Failure",
"steps": [
"1. Synthesia webhook → Make hook",
"2. Parse JSON for video URL",
"3. Router: If status = 'completed', proceed to Google Drive",
"4. Router: Else, wait 5 min, then re-poll API"
]
}
```

What else am I missing? Especially around rate limiting headers, data formatting for custom avatars/voices, or unexpected workflow bottlenecks you've hit.

— chloe


Webhooks or bust.


   
Quote
(@emma23)
Reputable Member
Joined: 2 months ago
Posts: 212
 

Great start with that list - those are the exact things that'll bite you later if you don't track them early.

I'd add **concurrent job handling**. Hit their API with 5-10 video requests at once during your trial and see if response times degrade or if you get throttling errors. That's where a lot of services show their true colors.

> connector quality for platforms like Zapier or Make
Test the lag time between Synthesia's webhook and the trigger in your automation. I've seen delays of several minutes that can mess up a workflow. Also, check if their native actions (like "create video") handle bulk operations or just one at a time.


Trial first, ask later.


   
ReplyQuote
(@heidir33)
Reputable Member
Joined: 2 months ago
Posts: 270
 

That's a really smart addition about concurrent jobs. In a live workflow, you almost never generate videos in a perfectly spaced queue, it's usually in bursts following a campaign launch or a segment update. I'm planning to simulate a realistic "Monday morning" spike scenario.

Your point about connector lag time is also crucial. When you mention "delays of several minutes," does that start to impact your retry logic or error handling? I'm thinking if the webhook is delayed beyond a certain window, our system might flag it as failed and retry, causing a duplicate. I'll need to build that buffer into my test.



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Your list is super solid for the technical side. I'd add one more category - human factors for your actual team's workflow.

What about the editor's usability for quick revisions? Time how long it takes to swap an avatar or fix a mispronounced word. If your content team isn't technical, a slow or clunky editor will kill adoption regardless of API reliability.

Also, track "time to final asset" from the first draft. Sometimes a 2-minute render is fast, but if you need 3 rounds of tweaks, the overall timeline balloons. That's the operational cost that doesn't show up in API logs.



   
ReplyQuote
(@infra_architect_rebel)
Honorable Member
Joined: 5 months ago
Posts: 544
 

Good list, but you're overcomplicating it. Skip the Zapier connector tests for now.

Focus on one thing: does the video arrive and work. Your entire workflow depends on that single event. All those other metrics are just symptoms if this fails.

Log the timestamp you call the API and the timestamp you have a usable file. That's your real SLA. Everything else is vendor KPIs they should handle.


Simplicity is the ultimate sophistication


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Totally agree on the concurrent jobs test, that's where the rubber meets the road. But for the Zapier/Make piece, I'd go one step further and say you should *only* test the native webhooks. If their own delivery is flaky, any connector built on top will be worse. The connectors are just another layer of potential failure you don't need.

Skip testing bulk operations in the connector. If their API can't handle a batch, the connector's just polishing a turd.


CRM is a necessary evil


   
ReplyQuote
(@data_analytics_rover)
Prominent Member
Joined: 6 months ago
Posts: 611
 

You're right to focus on the native webhook as the source of truth. The connector layer often introduces abstraction that just masks the underlying API's limitations.

I'd test the webhook's idempotency alongside its reliability. If a delayed delivery causes your system to retry, does the API treat the duplicate request gracefully or does it create a second, redundant video? That's a critical failure mode for automated workflows that's separate from simple uptime.

Testing the native batch endpoint directly gives you a clear performance benchmark. If it's slow or fails under load, you know the limitation is in their core infrastructure, not a third-party integration.



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

You're absolutely right that the human latency costs can dwarf API response times. I'd push one step further: log the *variability* in those editor task times. If swapping an avatar takes 2 seconds one day and 45 seconds the next, that's a huge red flag for team adoption. Unpredictable friction is worse than consistently slow.

Consider tracking "clicks to completion" for common edits too. An extra three clicks per fix, across dozens of videos a week, adds real operational drag the team will silently resent.



   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Good list, but you're missing the easiest one to fake in a trial: consistency in the generated video quality itself. Track the variance in lip-sync accuracy, background artifacts, and text-to-speech cadence across different avatars and scripts. A vendor can make one perfect demo video. Your team will make hundreds of slightly different ones, and that's where the uncanny valley will swallow your budget.

And while you're logging those render times, make sure you're doing it across different video lengths. A service that's optimized for 30-second social clips often falls apart on a 7-minute training module. The demo won't show you that.


cg


   
ReplyQuote
(@heatherm)
Reputable Member
Joined: 3 months ago
Posts: 255
 

Excellent point. The quality variance is the silent budget killer that only shows up at scale.

I'd add tracking consistency across *tiers of service* if they have them. Sometimes the "premium" avatar looks great, but the standard ones have noticeable lip-sync lag or robotic cadence. You need to know if the quality you're benchmarking is actually the one your team will use day-to-day.

Testing different video lengths is spot on. Also test with different languages or accents if that's in your scope. The TTS can handle a US English script perfectly and completely bungle a UK or Australian one.


Ask me about my RFP template


   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Exactly. You can't evaluate on their demo tier and buy the production tier. The performance characteristics and API rate limits often change between packages.

Set up your trial to match the exact subscription you'd actually purchase, not the free trial they give you. The free trial is often on premium infrastructure to hook you. Then you downgrade on commit and hit the real limits.

Also track how quality degrades under load. That premium avatar might look fine for your five test videos. Render fifty in a batch and see if the quality holds or if compression kicks in to save their compute costs.


Show me the query.


   
ReplyQuote
(@emilyr)
Reputable Member
Joined: 3 months ago
Posts: 295
 

Your list is a strong foundation. I'd add granularity to your API response times by separating out the p95 and p99 percentiles, not just averages. A service can have a great mean response time but still be unusable if 5% of requests stall for 30 seconds, which will break automated workflows. Also, categorize response times by the type of request: avatar generation vs. video generation vs. webhook callback, as they likely use different backend systems with distinct performance profiles.

Regarding **connector quality**, I'd advise against treating it as a primary evaluation metric. As others have hinted, it's a secondary layer. Your core test should be the native API's reliability and idempotency. If that's solid, a connector's failure is a solvable integration problem. If the API is flaky, no connector can save it. However, if you do test a connector like Zapier, log the added latency it introduces as a processing overhead cost, and more importantly, monitor for any data transformation errors it might inject between the webhook and your final destination.

The **cost per unit** metric is critical, but you need to define the "unit" rigorously. Is it per finished second of video, or per credit consumed? More importantly, track cost variance against the factors others mentioned: video length, avatar tier, and time of day. A price that scales linearly for a 30-second video but exponentially for a 10-minute one changes the business case entirely. Build a small script to graph cost against length and avatar type during your trial.



   
ReplyQuote