Skip to content
Notifications
Clear all

What's the fastest way to test 5 different system prompts in parallel?

4 Posts
4 Users
0 Reactions
29 Views
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
Topic starter   [#19711]

Hey everyone, I’ve been trying to improve some data quality checks by using an LLM to flag anomalies in our customer support logs. I have 5 different system prompt variations I want to test to see which one gives the most accurate and useful flags.

My current setup is a simple Python script that loops through the prompts and runs them sequentially against a sample dataset, but it’s super slow. I’m using the OpenAI API, and waiting for each one to finish before starting the next feels like watching paint dry 😅. I need to compare results across all 5 prompts on the same 1000 text snippets.

I’m thinking there must be a better way. I use Airflow for orchestration at work, but this feels like overkill for a one-off test. I’ve heard about using `asyncio` or maybe the `concurrent.futures` module, but I’m a bit lost on the best pattern. My constraints are:

* I need to log which prompt produced which response.
* I have a tight rate limit on the API (requests per minute).
* I’d like to avoid spinning up complex infrastructure.

What’s the simplest, most straightforward way to run these 5 prompts in parallel? Should I:
* Use `threading` or `asyncio` directly in a script?
* Write a small DAG that fires off 5 tasks?
* Is there a lightweight library built for this kind of prompt testing?

Any advice or examples of how you’ve done this would be a lifesaver. I’m worried about messing up the API rate limiting or losing track of the prompt-response mapping.


null


   
Quote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

I'd lean toward threading with a ThreadPoolExecutor for this one, mainly because the OpenAI Python client is synchronous under the hood. Asyncio can work if you wrap the calls properly, but for a quick one-off test, threading is simpler to reason about and you don't have to deal with an event loop.

The trick is your rate limit. You can pass a shared semaphore to each thread so you don't exceed the API's requests per minute. Something like a simple time-based token bucket in the worker function, or just a pause after every N calls. That way you can fire off all 5 prompts against the same 1000 snippets in parallel, but still respect the cap.

One thing to watch out for: logging which prompt produced which response. Make sure your worker function returns a tuple of (prompt_id, response) and then collate them after the pool completes. I've seen people accidentally lose that mapping when they write to a shared list without a lock.

Also, if your rate limit is really tight, you might get more throughput by sending the same prompt to multiple snippets in parallel rather than running all 5 prompts at once. That way you can saturate the limit per prompt without mixing up the results. Which approach are you leaning toward?



   
ReplyQuote
 dant
(@dant)
Honorable Member
Joined: 2 months ago
Posts: 434
 

Concurrent futures with ThreadPoolExecutor is indeed the pragmatic choice here, but I'd refine the approach on rate limiting. Wrapping synchronous HTTP calls in threads is straightforward, but a simple semaphore won't address the OpenAI tiered limits (requests per minute AND tokens per minute). You need client-side throttling that respects both.

Instead of baking logic into workers, use the `tenacity` library with a `wait_chain` for backoff on rate limit errors (status 429), and pair it with a global token bucket regulator for preventive throttling. This separates the retry logic from the concurrency mechanism. Also, for logging, don't just return a tuple - yield results incrementally to a thread-safe queue or write directly to a SQLite in-memory database with a connection per thread. This prevents memory bloat with 1000 snippets * 5 prompts.

One caveat: the official OpenAI client's internal session can cause contention under high concurrency. Consider creating a separate `requests.Session` or client instance per worker thread to avoid the GIL bottleneck on connection pooling.



   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

Since you're already leaning towards `concurrent.futures`, that's the right call for a quick script. The key is pairing it with the official OpenAI library's built-in retry logic. You don't need `tenacity` or a custom semaphore if you instantiate the client with `max_retries` and have a paid account; it handles 429s automatically.

Structure your script around a `ThreadPoolExecutor` with, say, 5 workers. Pass a `(prompt_id, prompt_text, snippet)` tuple to each submission. Have the worker function return that same tuple plus the response. Use `as_completed()` and write results immediately to a CSV. This gives you live logging and avoids memory issues with 1000 snippets.

Your main bottleneck will be tokens/minute, not requests. Threads will fire concurrently, so you'll hit that limit fast. A simple, crude fix is to batch your 1000 snippets into chunks of 20 and add a `time.sleep(60)` after each batch. It's not elegant, but for a one-off, it's simpler than implementing a token bucket.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote