Skip to content
Notifications
Clear all

Am I the only one who finds OpenAI's moderation API adds too much latency?

2 Posts
2 Users
0 Reactions
20 Views
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
Topic starter   [#14963]

I've been conducting a series of latency benchmarks for our text preprocessing pipeline, which includes a content safety check before sending prompts to a primary LLM. We've been testing OpenAI's moderation endpoint alongside a few other providers.

My initial findings indicate that the `moderations` API call adds a significant and consistent delay. In our controlled tests, using the `text-moderation-latest` model, the p95 latency added was between 120-180ms. This is on the same order of magnitude as the actual completion call to `gpt-4-turbo` for a short response.

This creates a tangible cost in terms of user-perceived latency, especially when it's a sequential, blocking step. For a high-volume application, this overhead compounds.

I'm curious if others have encountered this and what your strategies have been.
* Have you measured similar latency figures?
* Is anyone running the moderation check in parallel with other operations to mitigate the hit?
* Are there alternative content moderation services you've benchmarked that provide a better latency profile without sacrificing accuracy?

The reliability and accuracy are acceptable for our use case, but the performance cost seems high for a binary classification task.

—EK


Your bill is too high.


   
Quote
(@cloud_ops_amy_2)
Reputable Member
Joined: 7 months ago
Posts: 274
 

We saw similar latency in our setup. Your 120-180ms p95 aligns with our measurements last quarter.

We moved the moderation call to run in parallel with a non-blocking step, like fetching context from our vector DB. That cut the perceived latency almost to zero. It's not always possible depending on your flow, but worth checking if you can overlap any I/O.

For alternatives, AWS Bedrock's Content Moderation found a niche for us. The latency was lower on average, but you trade off for a slightly different classification taxonomy. You need to map their categories to your policy, which added some dev time.


terraform and chill


   
ReplyQuote