Skip to content
Notifications
Clear all

Why is Relevance AI so slow on large batch queries - any workarounds

2 Posts
2 Users
0 Reactions
18 Views
(@cloud_cost_watcher)
Honorable Member
Joined: 7 months ago
Posts: 386
Topic starter   [#25655]

Having integrated Relevance AI into our data enrichment pipeline for several months, I've observed a consistent and significant performance degradation when executing batch operations containing more than approximately 50-100 documents. While their vector search and data extraction capabilities are precise, the latency becomes a critical path blocker, directly impacting our cloud compute costs due to extended runtime.

The primary issue appears to be a lack of true asynchronous parallel processing at the API level. Even when we send a batch request, our monitoring suggests the tasks are being processed in a serialized or heavily throttled queue, rather than leveraging concurrent execution. This is particularly pronounced with workflows involving multiple AI models in sequence (e.g., extraction followed by summarization).

From a FinOps perspective, this forces an inefficient resource allocation. We must either:
* Accept prolonged Lambda or container runtimes, incurring higher compute charges.
* Implement complex client-side batching and retry logic, which adds development and maintenance overhead.
* Over-provision to meet time constraints, negating the cost-efficiency of a managed service.

I have attempted a few workarounds with limited success:
* Implementing an aggressive client-side sharding pattern, breaking one large batch into many smaller, concurrent API calls. This helps but hits other limits and multiplies network overhead.
* Experimenting with different `max_concurrency` settings within their workflow configuration, though the gains were marginal.
* Caching all possible intermediate results to avoid re-processing, which only mitigates the symptom.

My core question to the community is whether anyone has discovered a more effective architectural pattern or configuration setting to improve throughput. Specifically:
* Is there a documented way to leverage their async endpoints more effectively for true fire-and-forget batch processing?
* Has anyone benchmarked optimal batch sizes for different workflow complexities?
* Are there cost-effective complementary services (e.g., queue-based orchestrators) you've paired with Relevance AI to overcome this?

Optimize or die.


CloudCostHawk


   
Quote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

Yeah, they all throttle. Their "batch" endpoint is just a polite queue to manage their own infrastructure costs, not yours.

Your monitoring is right. I've seen this exact pattern with their workflow chains. Each step waits for the whole batch to finish before moving to the next model. It's serialized server-side, so your fancy async client code barely helps.

FinOps perspective is spot on though. The real cost is the dev time you burn building workarounds for their scaling limits. Sometimes the simpler, boring tool that does one thing fast is cheaper overall.


CRM is a means, not an end.


   
ReplyQuote