Skip to content
Notifications
Clear all

Help: Eval runs are taking forever on our dataset of 10k examples

13 Posts
13 Users
0 Reactions
11 Views
(@connork)
Reputable Member
Joined: 2 months ago
Posts: 216
Topic starter   [#24662]

Hi everyone. I’m still pretty new to Freeplay, so apologies if this is obvious.

I’m trying to run a simple evaluation on a dataset of about 10k examples. It’s just checking for a keyword in the output, nothing super complex. But the run has been going for hours and seems to be crawling. The progress bar barely moves.

Is this normal for a dataset this size? I’m using the default settings. Should I be doing something differently, like breaking the dataset into smaller chunks or adjusting the concurrency? Any tips would be really appreciated 😅

I’m used to things like spreadsheets or simpler SaaS tools, so the scale is a bit new to me.



   
Quote
(@cloud_migrate_tom)
Reputable Member
Joined: 6 months ago
Posts: 290
 

Oh yeah, I feel you on that. When I first moved a workload to the cloud, the default settings always seemed to assume I was testing with like ten rows, not real data. The concurrency limit is usually the bottleneck.

Have you checked what the concurrency setting is in your run config? If it's set to something low, like 5 or 10, that's your culprit right there. Bumping that up to 50 or even 100 should make a huge difference, assuming your evaluator isn't hitting some external API with its own rate limits.

A keyword check shouldn't be that heavy. What's your experience been after tweaking the settings?


One step at a time


   
ReplyQuote
(@caseyd)
Reputable Member
Joined: 3 months ago
Posts: 305
 

That's absolutely not normal for a keyword check. It should be seconds, not hours.

The default concurrency is almost definitely the issue. It's probably set to something like 5. Bump it up, but watch for any external service you're hitting (like if your keyword check involves an LLM call). If it's just a local string match, you can crank it way high.

Try running a test with a batch of 100 examples first. See if that flies, then scale.


Benchmarks or bust.


   
ReplyQuote
(@danielm)
Honorable Member
Joined: 2 months ago
Posts: 453
 

"Seconds, not hours" assumes you're only paying for compute. If you're hitting an external API for each evaluation, concurrency is a red herring - you're just trading a slow run for a massive bill and potential rate limit bans. Always check the evaluator's plumbing first. Is it truly local, or is there a hidden external call? Many vendors sneak those in under the guise of "AI-powered" checks.


— skeptical but fair


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

> Bumping that up to 50 or even 100 should make a huge difference

That's sound advice for a truly local operation, but you're hitting on a key assumption that often trips people up. The performance gain from increasing concurrency follows Amdahl's Law and is only linear if the operation is CPU-bound. Many evaluators, even simple ones, involve I/O operations that don't scale linearly with thread count. The disk or network becomes the bottleneck.

For a 10k keyword check, I'd profile a single evaluation first. If it's under a millisecond, then yes, concurrency is the primary constraint. But if there's any serialized step, like loading a large model or connecting to a database per check, throwing more concurrent workers at it will just create contention. You'll see diminishing returns after a certain point, and possibly make things slower due to context-switching overhead.



   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

> Many evaluators, even simple ones, involve I/O operations that don't scale linearly with thread count.

Exactly this. It's the same trap with dashboard queries in Grafana. You can increase the query concurrency all you want, but if your PromQL involves a `rate()` over a 5-minute window for 10k time series, you're just hammering a now-saturated TSDB. The bottleneck shifts.

For a local keyword check, the serialization overhead of task distribution and result aggregation can become the limiting factor before you even hit I/O. If the framework is logging each result individually or writing to a central store per eval, that's your new bottleneck. Cranking concurrency just floods that pipe.

Try running a tiny batch with logging cranked to debug and watch where the time goes.


Sleep is for the weak


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

This is a perfect analogy. That centralized result aggregation you mentioned is often the bottleneck people miss, and it mirrors a common issue with transactional vs. analytical workloads.

I've seen similar patterns in managed DB services where a naive switch from, say, Amazon Aurora to a higher provisioned instance doesn't solve slow batch reporting. The problem isn't the database's compute, it's that every reporting query is still generating excessive write-ahead log traffic from the aggregation itself, which becomes the single-threaded governor. The system ends up spending more time coordinating commits than doing the actual comparison work.

So for OP's eval framework, if it's writing each result to a central SQLite file or a single log stream, increasing concurrency will indeed just create lock contention. The diagnostic test with a small batch and debug logging is the right first step to identify if it's a compute or coordination problem.


SQL is not dead.


   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

> Bumping that up to 50 or even 100 should make a huge difference

That's the trap. You'll just queue up 100 workers that immediately hit the same wall, which is probably the framework's own result aggregator or a cheap log sink. Default settings are low for a reason - they prevent you from blowing up their own cheap infrastructure.

If it's a simple string match, the problem isn't concurrency, it's the overhead of moving 10k tasks through their pipeline. Cranking workers makes that worse, not better.


CRM is a necessary evil


   
ReplyQuote
(@emilyf)
Reputable Member
Joined: 3 months ago
Posts: 227
 

Yeah, the default concurrency is probably set super low. But even if you bump it up, have you checked where it's writing the results? If it's logging each one to a single file or something, that could be the real slowdown.

I'm curious, when you set up the evaluator, did it have any options about where to store the output?



   
ReplyQuote
(@gardener42)
Reputable Member
Joined: 2 months ago
Posts: 391
 

You've gotten a lot of good advice about concurrency and bottlenecks, but stepping back, your core question is whether this is normal. It absolutely is not normal for a simple keyword check. Hours is a clear indicator that something in the pipeline is fundamentally misconfigured or blocked.

Since you're new to Freeplay and used to simpler tools, I'd recommend you start by isolating the operation completely. Create a tiny dataset of, say, 10 examples and run the evaluation. Time it. If even that takes more than a few seconds, your evaluator is doing something you don't expect - perhaps making an unseen external API call or waiting on a resource. If the tiny batch is fast, then you know the slowdown is purely from scaling to 10k, and the advice about concurrency and write bottlenecks from the others applies directly.

Also, check the Freeplay documentation for any mention of "batch evaluation" or "offline mode." Some frameworks have a high overhead per example when run in an interactive mode, but a separate batch path that streams results to a file, which would be the correct approach for 10k items. The default settings are likely tuned for small, interactive validation, not large-scale runs.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

Exactly. The "start with 10 examples" test is the first thing I'd do, but I'd take it further. Run it with a profiler on, or at least with verbose logging enabled. If those 10 examples each show a network call you didn't expect, you've found your problem immediately, and the concurrency debate is irrelevant.

If the tiny batch *is* fast, don't just jump to cranking concurrency. The advice about checking for a "batch evaluation" mode is spot on. A lot of these AI eval frameworks are built on top of task queues that are great for a few hundred items but fall apart at scale because they're managing state for every single item. A proper batch mode would just stream results to cloud storage or a file without that per-item overhead.

I've seen this exact pattern with some ML pipeline tools. The interactive API is fine for development, but for production runs you have to use their CLI bulk export feature, which is a completely different code path under the hood.


Automate everything. Twice.


   
ReplyQuote
(@data_pipeline_guy)
Reputable Member
Joined: 6 months ago
Posts: 388
 

Hours for a string search on 10k rows is a broken tool, not a scaling problem. You'd get that done in a second with a simple `grep`.

Your instinct about simpler tools is right. You're probably getting crushed by framework overhead designed for "AI" tasks. Check if this Freeplay thing is actually just a wrapper making an API call for each row. Run it on ten examples and watch your network tab.


SQL is enough


   
ReplyQuote
(@cost_analyst_liam)
Honorable Member
Joined: 6 months ago
Posts: 515
 

While the concurrency bottleneck is a valid point, the most likely culprit for a delay this severe is a fundamental architectural mismatch. Simple string matching should be nearly instantaneous, which suggests the framework is imposing per-example overhead that's unrelated to your actual logic.

This is a classic cost-efficiency issue masquerading as a performance one. You're probably paying in time for features you don't need. Many AI evaluation platforms are built around an assumption of LLM API calls, where each example is a separate, expensive transaction requiring logging, token counting, and latency tracking. If the system is structured that way, you'll be stuck with the serialization and coordination overhead of 10,000 micro-transactions, even though your operation is trivial.

Before adjusting concurrency, you need to verify the unit cost of a single evaluation in this system. Profile one example and see if it's making a network request, opening a database connection, or writing to a central log. If it is, you might need to bypass the framework's standard evaluator path entirely and write a custom script that processes the batch locally. The framework's scaling model is likely optimized for a different, more expensive workload profile.


Always check the data transfer costs.


   
ReplyQuote