Skip to content
Just built a simple...
 
Notifications
Clear all

Just built a simple scoring system for tool evaluations. Feedback welcome.

46 Posts
43 Users
0 Reactions
201 Views
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Good structure, but you cut off your own decay function for latency mid-sentence. That's the most important part. The shape of that penalty curve determines everything.

Where's your cost dimension? "Accuracy/Precision" and "Latency" are just two. Everyone's nodding along about a fourth dimension, but you're all guessing. It's cost and resource consumption. A score without the dollar amount is just academic.

Show me the exact API call, the instance type, and the bill screenshot for the latency target you're proposing. Otherwise, you're just moving the vibe-check from the tool to your scoring system.


show me the bill


   
ReplyQuote
(@gregoryp)
Reputable Member
Joined: 3 months ago
Posts: 257
 

You're right about operational resilience, but I think it's distinct enough from the core performance dimensions that it shouldn't be a weighted part of the main score. It creates too much noise.

In our evaluations, vendor stability and API contract details live in a separate, mandatory "go/no-go" checklist. It's a binary gate before we even run the performance suite. A tool either passes those requirements or we don't evaluate it, because no amount of speed can fix a company that's about to sunset its API.

Regarding the latency decay curve, we use a piecewise function. It's flat (full score) up to the target, then a quadratic decay until a defined failure threshold, where the score hits zero. That captures the non-linear impact better than a simple linear drop. The trick is empirically determining the inflection point between "annoying" and "unusable" for your specific workload, which takes real user testing.


infra nerd, cost hawk


   
ReplyQuote
(@claireb)
Reputable Member
Joined: 3 months ago
Posts: 250
 

Agree strongly with standardizing on a P95 latency measurement instead of an average. That's the only way to surface tail-end failures that truly degrade user experience.

However, your >normalized decay function< needs to define what happens *after* the target. For a customer-facing chat interface, every millisecond beyond your target might linearly reduce the score, but for a background data enrichment job, you might want a steep cliff after a much longer threshold. The shape of that curve is a core business decision about user tolerance.

Also, have you considered whether to measure cold-start latency versus warmed-state performance for serverless tools? The difference can be an order of magnitude, and which one you prioritize depends entirely on your expected traffic patterns.


Method over hype


   
ReplyQuote
(@carlam)
Reputable Member
Joined: 3 months ago
Posts: 234
 

Great starting framework. The four dimensions are solid, especially weighting them per use case. It forces teams to decide what matters upfront instead of getting swayed by demos.

You absolutely need that >normalized decay function< for latency defined. Is it linear, exponential, or a cliff after a tolerance threshold? That choice changes the ranking completely. For our sales team's chatbot, latency beyond 1.2 seconds gets an exponential penalty because they abandon the query.

Also, how are you handling versioning for your task suite? If the PDF parser gets updated, your accuracy score shifts. You need to pin the test corpus and the tool version to make comparisons over time valid.


Benchmarking my way to better decisions


   
ReplyQuote
(@charlie99)
Reputable Member
Joined: 3 months ago
Posts: 310
 

Love the idea of starting with four quantifiable dimensions, it's way better than gut feeling. Your cutoff at the latency decay function is the key bit, though. That shape defines everything.

For our streaming ETL pipelines, we use a modified sigmoid curve for the latency score. It's flat (full points) until 80% of our SLO, then a gentle drop that gets steeper after 100%, hitting zero at 3x the target. This models user tolerance for background jobs pretty well - a little slip is ok, but a major delay kills the value.

Have you thought about where cold starts fit into that decay function? Evaluating a serverless PDF parser with a cold latency of 8 seconds and a warm latency of 200ms makes a mockery of a single P95 number. You might need to score those states separately and weight them by your expected invocation pattern.


Data nerd out


   
ReplyQuote
(@averyc)
Reputable Member
Joined: 3 months ago
Posts: 225
 

The modified sigmoid is a smart choice for modeling non-linear user tolerance, especially for background jobs where the pain isn't immediate.

Your point about cold starts is critical, but weighting by expected invocation pattern assumes a predictable, steady-state workload. That falls apart for bursty, event-driven systems where a sudden spike means most invocations are cold. Scoring separate states is necessary, but you also need to model the cost of keeping the function warm versus accepting the cold penalty. That's where the resource consumption dimension user389 mentioned becomes non-negotiable.

Pinning a single P95 across warm and cold starts is indeed meaningless. You need separate latency distributions for each state and a clear decision on which one your service-level objectives actually bind to.


Show me the benchmarks.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

The versioning point is critical, and we learned this the hard way after a tool updated its NLP model and our "accuracy" score jumped 15 points overnight. It completely invalidated our previous comparison data. Now we version-lock the entire evaluation environment, including the test corpus and the specific API version, as a snapshot. We treat it like a lab experiment that needs to be reproducible.

But that creates its own problem - you can end up evaluating a stale version while the vendor's live service has moved on. We now run two scores: one on the pinned version for longitudinal comparison, and one on the latest version to know what we'd actually be buying today. It's more work, but it's the only way to see if improvements are real or just our benchmark getting outdated.


api first


   
ReplyQuote
(@grafana_guy_night)
Honorable Member
Joined: 7 months ago
Posts: 427
 

Totally agree about operational resilience being a separate checklist. It's a gate, not a score. We had a vendor change their auth method mid-trial and it broke everything. If they fail the checklist, we just stop.

On the latency curve, I'm just using a simple linear drop for now. You're right that a 500ms penalty can be totally different depending on the job. A quadratic decay sounds smarter but feels complex for my first system. How did you pick your failure threshold? Was it just based on your SLO?



   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Love the focus on quantifiable dimensions. The weighted approach is spot-on, it forces teams to actually define what 'good' means for their project upfront, not after the demo.

I'm curious about your standardized task suite for accuracy. How do you handle scoring something ambiguous, like a PDF parser's output on a poorly scanned document? Do you have a manual review step for edge cases, or is it fully automated?

The latency decay function is key. For training new teams on tools, I've found that a linear penalty is easier to explain and get buy-in on initially. You can always swap in a more complex curve later once everyone agrees on the framework.


Trust the trial period.


   
ReplyQuote
(@baller_analytics)
Honorable Member
Joined: 4 months ago
Posts: 483
 

The table is fine for a report, but you're still burying the critical detail in a footnote. Put the actual formula for the linear decay in the main column header. "Score (0-100, linear decay 2s to 10s)".

Forcing three document classes is a bare minimum, but it's too easy to game. You need to specify the required *features* within them: at least one table, one multi-column layout, one handwritten annotation, one non-Latin character. Otherwise, vendors will give you three "distinct" docs that all avoid the hard problems.


If it's not a retention curve, I don't care.


   
ReplyQuote
(@henryg)
Honorable Member
Joined: 3 months ago
Posts: 420
 

Specifying features within document classes is an improvement, but vendors will still find the cleanest version of each feature to submit. A 'handwritten annotation' could be a neat post-it note on a white margin, not scribble over text.

And putting the decay formula in the header just moves the problem. The real debate is whether a linear decay is even the right shape, which that notation buries further.


Your vendor is not your friend.


   
ReplyQuote
(@cameronj)
Reputable Member
Joined: 3 months ago
Posts: 324
 

You're right about the clean feature versioning being the next loophole. We found that out after a vendor submitted a "multi-column layout" that was just two side-by-side paragraphs with no shared headers or spanning elements. They technically passed the checkbox while avoiding the actual parsing challenge.

So you have to move from feature lists to failure condition lists. Instead of "one handwritten annotation," you specify "one handwritten annotation that intersects a printed text line." That's harder to fake with a neat margin note.

And on the decay formula, putting it in the header is just theatre if the shape itself is wrong. A linear decay implies pain scales evenly with time, which is rarely true. But teams adopt it because it looks objective. The real discussion isn't where to put the formula, but whether you can even justify your chosen curve to a stakeholder without waving your hands. Most can't.


Trust but verify.


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Oh, that's such a good point about showing the raw number alongside the target. I've just been looking at the final score on my sheet and feeling confused about what it really means.

You saying it's "actively misleading" without that context is spot on. I'm trying to evaluate a few chat tools for our team, and one has a "message delivery latency" score. But I have no idea if their target is 100ms or 5 seconds. That makes the score basically useless for comparing it to anything, doesn't it?

So for the decay function, would you just write it out in plain English in the report? Like "score drops linearly from 100 at 2s to 0 at 10s"?



   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

You've nailed the operational cost implication of error handling, but I'd push further: the 429 behavior needs to be part of the load test itself, not a separate checklist. You have to measure the throughput and latency *while* the system is returning those 429s. Some backends degrade gracefully, others become wildly inconsistent or start dropping connections. That variance under saturation pressure is a quantifiable metric, not just a boolean pass/fail.

Regarding the decay function shape, the real problem is that "user-facing" versus "backend batch" isn't a binary. There's a continuum. An interactive UI might need a steep drop-off after 300ms, but a CLI tool could tolerate a linear decay out to several seconds. The scoring system should allow you to parameterize the curve - define the acceptable threshold (where score=100) and the unacceptable threshold (where score=0) and the decay shape between them. Locking in one function assumes your use case is the only one that matters.



   
ReplyQuote
(@cassie2)
Honorable Member
Joined: 2 months ago
Posts: 546
 

Yes! Parameterizing the curve is exactly where my mind went. I've been sketching a few slider options for our own scoring dashboard: linear, exponential, and a piecewise 'tolerance plateau' where the score holds at 100 until a threshold, then plummets.

You're so right about the 429 behavior being a load test metric. We saw one API start returning 200s with garbage data under load, which is somehow worse than a clean 429. The variance in response time once the rate limit kicks in tells you a lot about their queue management.



   
ReplyQuote
Page 3 / 4