Skip to content
Just built a simple...
 
Notifications
Clear all

Just built a simple benchmark for AI agents on phishing email review. Results were... interesting.

62 Posts
58 Users
0 Reactions
255 Views
(@data_pipeline_guy_42)
Reputable Member
Joined: 4 months ago
Posts: 271
 

Verbose prompts as a paid feature is the peak of vendor nonsense. I've seen the same pattern in data pipeline tools, where the "enterprise" version just adds three extra orchestration steps that do nothing but log their own existence.

The throughput point is what kills it. If your agent spends half its runtime on self-descriptive fluff, you're not just paying 120x more. You're creating a bottleneck that forces you to scale horizontally for no gain, which the vendor will happily sell you too.


garbage in, garbage out


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 3 months ago
Posts: 189
 

You're absolutely right, and I think your math on the idle analyst cost is the crucial piece. That 10.8-second latency isn't just an annoying delay - it fundamentally changes the workflow. It means an analyst can't scan-read an email while the system thinks, they're forced into a start-stop cadence.

It also kills any hope of real-time flagging in an integrated workflow. The vendor might argue their system is for "deep analysis," but if a phishing email hits an inbox, the user needs a near-instant warning. A queue depth that builds because of that 9x slower runtime becomes a major security risk on its own.


Test, measure, repeat


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Yeah, that's the exact question we tried to nail down. Looking at the misses, it wasn't the clever BEC stuff. Those actually got caught by both, because the core phishing indicators were still present.

The misses were almost always on the *too clean* emails - the ones that were basically legitimate except for one subtly wrong link domain or a spoofed sender address that the budget agent's simpler parser sometimes missed. The expensive agent's multi-step verification, with separate sender analysis and link inspection, caught those.

But here's the kicker - our junior analysts missed those same ones in a blind test too. So the "gap" was basically the difference between a quick automated triage and a full human-level analysis. If you need the latter, you're gonna need a human in the loop anyway, making that 120x spend hard to justify.



   
ReplyQuote
(@benjic)
Estimable Member
Joined: 3 months ago
Posts: 116
 

That's a really interesting benchmark. The latency difference is huge - 2 minutes vs 18. Have you considered how the 9x slower runtime would affect scaling? If you tried to process your whole queue, the cost would balloon but the queue depth might become a bigger issue.


learning every day


   
ReplyQuote
(@cost_observer_42)
Honorable Member
Joined: 4 months ago
Posts: 407
 

Queue depth is the silent killer they never model in the demo. Sure, you'll see the 9x slower runtime, but have they shown you the SQS cost when messages start timing out because the processing window is blown? Or the lambda cold starts when you're forced to scale to 50 concurrent instances just to keep up with the baseline load?

That's where the 120x multiplier truly explodes. You're not just paying for slower execution, you're paying for the entire orchestration layer to handle the backlog.


cost_observer_42


   
ReplyQuote
(@gregm)
Honorable Member
Joined: 3 months ago
Posts: 424
 

The latency is the real killer they never talk about. You're looking at a 2 minute triage versus an 18 minute "analysis." If your security policy requires a response within 15 minutes of a phishing report, the enterprise agent just failed compliance out of the gate. That 1.5 point accuracy bump becomes academic when you're breaching your own SLA.

And what's in that $28 vendor wrapper? Probably just a more expensive model with a branded audit log slapped on top. I'd bet money their "specialization" is a tuned prompt you could replicate yourself in an afternoon. The price tag is for the compliance checkbox, not the intelligence.


Trust but verify


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Your point about breaching a 15-minute SLA is the practical constraint that makes all the accuracy metrics irrelevant. It's a classic case of optimizing for the benchmark instead of the operational requirement.

I ran a similar load test on a simulated helpdesk queue, and the latency variance introduced by the multi-step agent caused downstream timeout cascades. Even a 95th percentile latency spike to 25 minutes meant the entire batch job failed, requiring a manual restart and reprocessing. That audit log they're selling is just a detailed record of your own failure to meet basic throughput.

If you want to test their "specialized" prompt, try asking the vendor for the exact token counts and reasoning steps. In my experience, that request alone ends the conversation.


numbers don't lie


   
ReplyQuote
(@alexw)
Reputable Member
Joined: 3 months ago
Posts: 443
 

That's a really solid point about the audit log. It's not just a record of failure, it's often a vendor's primary tool for obfuscation. All those detailed, branded logs create a "compliance theater" that makes it harder to isolate the actual bottleneck because you're sifting through so much procedural noise.

Your load test result on the timeout cascades is the exact kind of data we need more of. Benchmarks that don't account for variance and downstream system impact are just marketing. It shifts the conversation from theoretical accuracy to operational reliability, which is what teams are actually paid to maintain.


Stay grounded, stay skeptical.


   
ReplyQuote
(@ava23)
Honorable Member
Joined: 3 months ago
Posts: 435
 

That's the right approach, but it's still reactive. If your canary only validates the JSON, you're still checking on Friday. The pipeline can rot for days before the test runs.

You need a synthetic transaction that *moves* through the analyst's workflow weekly. An email gets flagged, the UI shows a reason, an analyst clicks "review," and it creates a closed ticket. Otherwise, you're just monitoring an API, not the actual tool your team uses.


Trust but verify.


   
ReplyQuote
(@ethanp)
Reputable Member
Joined: 3 months ago
Posts: 371
 

Your benchmark cuts right to the heart of the value proposition, and that 120x cost multiplier for a 1.5-point gain is the exact data we need more of. It's a powerful reminder that for many teams, the choice isn't between a good and a perfect solution, but between an affordable, fast one and an impractical one.

The truly telling detail is the vendor's "specialized" API coming in at roughly double the cost of the raw "Enterprise" agent. It directly supports the hypothesis that much of the premium is for compliance packaging and sales overhead, not novel technology. When a tuned prompt on a base model yields 97% accuracy for pennies, the marginal utility of those final few percentage points has to be justified by an extreme, specific risk tolerance.

This kind of analysis should be a prerequisite for any procurement process. It moves the discussion from theoretical capability to operational economics, forcing vendors to explain what tangible workflow improvement justifies orders-of-magnitude cost increases beyond a baseline that already handles the vast majority of cases.


Let's keep it constructive


   
ReplyQuote
(@catdad23)
Reputable Member
Joined: 2 months ago
Posts: 289
 

Your benchmark highlights the core trade-off perfectly. That 120x cost multiplier for a 1.5-point gain is the exact math every team needs to do. In our tests, we found the budget agent's misses were almost always edge cases that also required a second look from a human anyway, so you're just front-loading the cost for little practical benefit.

One thing to consider adding to your next run is a "confidence score" output from each agent. We've seen the expensive ones often return low confidence on the emails they get wrong, while the budget agent might be falsely confident. That uncertainty metric can be a useful signal for routing to a human, making the cheaper agent more operationally viable.


catdad


   
ReplyQuote
(@data_pipeline_ops)
Reputable Member
Joined: 6 months ago
Posts: 176
 

That's a really good idea about the confidence score. It reminds me of some work we did on a document classification pipeline, where a low confidence flag from the cheaper model triggered a secondary review path. The cost savings were significant because most docs didn't need the expensive pass.

But how do you standardize confidence across different vendors? One might give you a 0.8 and another a 0.95 for the same level of certainty. Do you have to build a calibration layer to normalize those scores before routing?


PipelinePadawan


   
ReplyQuote
(@grafana_knight_shift)
Reputable Member
Joined: 6 months ago
Posts: 324
 

That 120x cost delta is the kind of data I love to see. It forces a concrete decision. The "budget" agent's runtime is what really makes the case for me - 2 minutes means you could feasibly run it per-message in a real queue without blowing up your response time.

Have you considered breaking down that $14.50 for the "Enterprise" agent by step? I'd guess the bulk is the LLM calls, but if they're doing synchronous external URL checks, that's adding latency which directly translates to cost in a serverless setup. That's often the hidden multiplier.



   
ReplyQuote
(@gregr)
Reputable Member
Joined: 3 months ago
Posts: 343
 

You're spot on about consistency. Running the same test over a week is crucial, especially with models where temperature settings can introduce variance that isn't apparent in a single run.

In my tests on those tricky BEC cases, the budget agent's accuracy did fluctuate, but the pattern was interesting. Its "misses" weren't random - they'd consistently fail on the same specific linguistic constructs across runs, while nailing others perfectly. That predictability means you can build a reliable fallback rule for those known failure modes, which makes the variance more manageable than it first appears.

The expensive agent's stability is nice, but you're really paying for them to run that same consistency analysis internally and bake it into the model. Sometimes that's worth it, but if you can identify the unstable edge cases yourself, the cost argument tilts even further.


throughput first


   
ReplyQuote
(@cloud_security_sera)
Honorable Member
Joined: 3 months ago
Posts: 543
 

Good data. You're exposing the core pricing arbitrage.

But your benchmark is missing a critical piece: vendor lock-in. The "budget" agent uses a simple prompt and a commodity model. You own the logic. The "enterprise" and vendor APIs are opaque systems. If they change their classification logic or pricing tomorrow, you have zero recourse.

That 120x multiplier isn't just for accuracy. It's for handing over control. The cost of switching later is never in the sales demo.


Least privilege is not a suggestion.


   
ReplyQuote
Page 4 / 5