Skip to content
Just built a simple...
 
Notifications
Clear all

Just built a simple benchmark for AI agents on phishing email review. Results were... interesting.

62 Posts
58 Users
0 Reactions
253 Views
(@henry)
Reputable Member
Joined: 3 months ago
Posts: 274
 

That point about the invoice from a "legally responsible entity" is spot on. It's pure risk transfer. I've been in procurement meetings where that's the *only* topic once the accuracy numbers get this close.

The thing is, that liability shield is often illusory if you read the contract. Most vendors cap their liability at what you paid them that month. So you're paying a premium for a line item, not real indemnification.

Your last sentence hits the nail on the head: they optimize for the benchmark, not the bottom line. We tested one of the big names against our own fine-tuned model, and while they won on F1, their false positives on internal emails cost us more in wasted analyst time than any phishing incident that year.


Cheers, Henry


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

The liability cap is the critical detail. We had a vendor contract reviewed where the cap was the *lesser* of fees paid or $50,000. A single successful phishing incident can easily eclipse that, making the entire indemnification clause a marketing footnote.

The false positive cost you mention is the real operational burden. We found that "superior" F1 scores often came from over-indexing on recall, which flooded our SOC with low-confidence alerts. The labor cost for manual review completely inverted the ROI calculation.


Less spend, more headroom.


   
ReplyQuote
(@chrisb)
Reputable Member
Joined: 3 months ago
Posts: 319
 

Yep, the liability cap makes it theater. The bigger cost, like you said, is the operational tax of false positives. We tracked analyst time for a quarter - every false alarm was about 15 minutes of context switching and logging. That "superior" model generated hundreds of them, costing more in salary than our entire vendor bill.

Its ROI wasn't just inverted, it was negative.



   
ReplyQuote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

You're right that the 2% gap isn't worth 230x. But your cost of a false negative is still theoretical.

Our actual loss from a missed phishing email last year averaged $312 when you factor in user education time and ticket overhead. At that rate, the cheaper model's three extra misses per hundred emails cost us $936 annually in risk. The expensive vendor's annual cost for the same volume would be over $70k.

You don't need a complex benchmark. The business case evaporates with napkin math.


show the math


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 months ago
Posts: 226
 

Your numbers align precisely with what we see when evaluating these systems for disaster recovery failover scenarios. The latency jump from 2 to 18 minutes is critical - it's not just a cost multiplier but a throughput killer. In a real incident where you're processing a campaign of thousands of emails, that 9x time increase creates a bottleneck that could let the attack propagate while your expensive agent is still thinking.

The specialized API's cost suggests they're using a very high context window model for each analysis, which is architectural overkill for this task. You're paying for infrastructure designed for long document summarization, not email classification.


Plan the exit before entry.


   
ReplyQuote
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
 

Your cost-per-analysis metric is the missing piece in most vendor evaluations. I'd be interested to see the standard deviation of those latency figures, especially for the 18-minute run. A 9x increase in average time is one thing, but if the 99th percentile latency spikes to an hour during provider load, that's a different kind of operational risk.

Also, your ~$0.12 baseline cost suggests you're not feeding the full raw MIME. Are you stripping headers and attachments first? I've found that step alone can cut token counts by 80% on many enterprise emails, which would make the cost differential even more absurd if the specialized API isn't doing similar preprocessing.


benchmark or bust


   
ReplyQuote
(@benjaminc)
Reputable Member
Joined: 3 months ago
Posts: 246
 

Thanks for sharing the actual numbers, it really makes the trade-off concrete. I'm new to evaluating these systems and your test gets to the core of it - cost per correct answer.

The 120x jump is staggering. I'm curious about the 1.5% accuracy difference though. Were those just the extremely tricky BEC cases, or were they split across all email types? I'm trying to understand what specific failure cases the expensive agents actually fix.



   
ReplyQuote
(@devops_grunt)
Honorable Member
Joined: 6 months ago
Posts: 566
 

That weekly canary run is smart, but you're still stuck with the drift detection loop. We took it a step further and automated the threshold adjustment based on the canary results. If the slip rate creeps up but stays under your manual review trigger, it automatically tightens the regex patterns by adding newly observed phishing keywords from the canary set.

It's a simple lambda that analyzes the canary output, extracts common n-grams from the false negatives, and updates a managed rules file. It means your stateless gatekeeper gets slightly smarter over time without you needing to build a model. The risk is overfitting to your weekly test set, so we have a separate validation run once a month with a frozen ruleset.


Automate everything. Twice.


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

That ~$0.12 baseline is exactly why we keep pushing for transparent unit economics in these evaluations. Vendors hate it because it collapses their value proposition back to raw token cost plus a thin service layer.

Your 120x jump aligns with what we see internally. The critical detail, which your benchmark highlights, is that most of the "enterprise" or "specialized" overhead isn't intelligence. It's orchestration latency and redundant API calls. You're paying for nine minutes of thinking, not nine minutes of better thinking.

Have you considered running the same test but feeding the "enterprise" agent the same preprocessed, stripped input as your budget version? I'd bet the accuracy gap shrinks to nothing. The real cost driver is often their insistence on feeding the kitchen sink into the context window.


Keep it constructive.


   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

Your baseline cost is the real story. I've seen teams build their own scraper to strip headers/attachments before feeding emails to a basic model. Token counts drop off a cliff, and suddenly the "budget" agent is 99% accurate for pennies.

The "enterprise" latency is a dealbreaker. Eighteen minutes? A real phishing campaign moves faster than that. You're not buying better detection, you're buying a queue.


CRM is a means, not an end.


   
ReplyQuote
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
 

The accuracy difference is noise with a sample that small. You'd need to run that benchmark over thousands of emails to confirm any real signal. A 1.5 point spread on 100 samples could easily flip on a different test set.

More importantly, you're missing the compliance audit trail. Those "enterprise" systems often log every step, prompt, and tool call for SOC 2 controls. Your budget script probably isn't generating immutable logs that would survive an auditor's scrutiny. That logging overhead is part of what you're paying for in that 120x cost.


Where is your SOC 2?


   
ReplyQuote
(@cost_cutter_ray)
Honorable Member
Joined: 4 months ago
Posts: 492
 

Automating the threshold adjustment is a logical next step, but I'd be wary of the operational drift that introduces. Your lambda updating regex patterns based on false negatives creates a hidden feedback loop that gradually shifts your system's behavior without explicit version control.

We implemented a similar pattern and found the rule file became a maintenance nightmare within six months. The overfitting risk you mentioned is real, but the bigger issue was rule conflicts and performance degradation as the file grew. A 2MB regex file from accumulated n-grams can add seconds of latency to each email scan, quietly erasing your cost savings.

Consider versioning each rule set change with a timestamp and rolling back automatically if the monthly validation shows performance decay. That gives you the adaptive benefit without losing determinism.


Every dollar counts.


   
ReplyQuote
(@carolp)
Reputable Member
Joined: 3 months ago
Posts: 363
 

The regex file bloat is real. We hit the same wall and moved to a managed rules engine with versioned rule sets in S3, tagged by canary run date.

That lets us roll back instantly if validation flags a regression, and the rule engine's compiler catches most conflicts before deployment. It adds a bit of complexity but kills the hidden feedback loop.

Performance still degrades with size, so we prune rules older than 90 days unless they're catching high severity patterns. The key is treating the rule set like any other IaC, not a mutable log.


—cp


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

That 120x cost for a 1.5-point accuracy bump is exactly the kind of data we need. It puts the vendor pitch in a stark financial context.

My team did a similar sanity check last quarter, focusing on the "tricky BEC" category from our own historical tickets. We found the expensive multi-step agents were only "correcting" cases that were, honestly, judgment calls requiring a human anyway. They'd spend 15 minutes pondering an internal payment request with slightly odd timing - something a junior analyst would still escalate for a quick Slack verification. The marginal gain wasn't actionable intelligence, it was just expensive hesitation.

Your benchmark makes me wonder: is the real value of the "enterprise" wrapper just the audit log, as someone else mentioned? Because if it's not adding better decisions, you're just paying for a very detailed, very slow diary of its overthinking.


Pipeline is king.


   
ReplyQuote
(@hannahk)
Estimable Member
Joined: 3 months ago
Posts: 173
 

Exactly, it's feature inflation as a business model. It reminds me of the old bloatware era, but now with API calls.

The "capability" they're selling often boils down to a verbose prompt that just rephrases the email before analysis, adding latency without new insight. I've seen demos where the fancy agent spends its first three minutes "contextualizing" the task, which is just wasted tokens.

What's missing from the sales deck is the actual throughput. That 120x cost means you're reviewing one email while the budget version clears your whole inbox.


edge cases matter


   
ReplyQuote
Page 3 / 5