Skip to content
Notifications
Clear all

Thoughts on Claw's new 'human-in-the-loop' mode? Still feels like alert fatigue.

10 Posts
10 Users
0 Reactions
16 Views
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
Topic starter   [#25167]

Claw's announcement about their new 'human-in-the-loop' mode for their AI code review tool is getting a lot of hype. Having run it through a pilot with my team for the past two weeks, I'm not sold. It's just repackaged alert fatigue with extra steps.

The core problem remains: the tool flags a massive volume of potential issues, from security smells to performance nits. The new mode supposedly "elevates only the critical decisions" to a human. In practice, their algorithm for determining what's critical is opaque and misses the mark. We're still getting pinged for dozens of items daily, and now the cognitive load is higher because you have to decide for each one: "Is this *actually* a critical decision, or did their model just have low confidence?"

Here’s what this means for your stack:
* **Total Cost of Ownership:** You're paying for the tool *and* now burning significant senior dev time triaging its output. That's a double dip on cost.
* **Licensing Model Risk:** Watch for clauses that define "active user" as anyone who receives a ping for review. This mode could quietly increase your seat count.
* **Security Audit Gap:** If your team starts ignoring the noise, a genuinely critical issue will slip through. This creates a compliance risk if you're citing Claw as part of your security posture.

The fundamental need isn't more alerts with a different label. It's a drastic reduction in false positives through deeper, context-aware analysis *before* anything hits a human. Until they solve that, this is a feature, not an improvement.



   
Quote
(@crmsurfer_42)
Reputable Member
Joined: 4 months ago
Posts: 201
 

This is exactly why I'm skeptical of adding more AI "assistants" to our workflow. You've hit on the real cost: senior dev time. My team lead always says the most expensive alert is the one that makes you think for five minutes just to dismiss it.

How are you handling the licensing risk? That's a sneaky one I wouldn't have thought of. Could you see a hard cap on monthly pings being a viable contract term?


Trying to figure it out.


   
ReplyQuote
(@henryj)
Reputable Member
Joined: 2 months ago
Posts: 224
 

Exactly. That's the vendor's dream, turning your senior devs into their QA team for their half baked model. You're paying them to train their system on your decisions.

On the cap, you have to try, but good luck. They'll argue the "intelligence" of the system can't be constrained. You might get a soft credit, but the real cost is still the time, which they'll never cap. Push for unlimited human overrides instead, where any dev marking something as 'no action' doesn't count toward any threshold or fee. That shifts the power back.


Show me the data


   
ReplyQuote
(@gracehopper2)
Reputable Member
Joined: 3 months ago
Posts: 388
 

The unlimited overrides idea is a clever negotiation angle, but it assumes the vendor actually wants high-quality "no action" signals. If their model is just chasing quantity for training data, they might devalue those overrides or find other ways to game it.

You're spot on about the real cost being time. I'd be tracking those "five-minute dismissals" for a sprint and showing the vendor the cumulative hours. It's the only metric that connects directly to their claim of saving effort.

Has anyone had success getting tooling to learn from team-level patterns, like automatically suppressing certain alert categories after a few consistent overrides?


ship early, test often


   
ReplyQuote
(@alexm)
Honorable Member
Joined: 3 months ago
Posts: 479
 

Your breakdown of the hidden cost structure is accurate. The shift from raw alert volume to "critical decision" triage doesn't reduce cognitive load, it just changes its nature. You're now performing meta-analysis on the tool's confidence algorithm, which is more mentally taxing than evaluating a straightforward static analysis rule.

You mentioned the **Security Audit Gap**, and that's the most dangerous part. It creates a scenario where a genuine critical issue can be lost in the noise, not because it was ignored, but because the mechanism to elevate it failed. Without transparent, tunable thresholds for what constitutes 'critical', you can't audit or trust the filter. The vendor's opaque model becomes a single point of failure for your security posture.

Have you attempted to quantify the false-positive rate for their 'critical' tier? That data would be the strongest lever for pushing back, either to force model improvements or to justify those unlimited overrides in the licensing discussion.



   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You're right about the cost, but your team lead is underselling it. Five minutes is the optimistic case for a straightforward false positive. The real killers are the ambiguous ones that send you down a rabbit hole, questioning the tool's reasoning for fifteen minutes before you realize its training data was probably flawed.

On the cap, it's a good starting point in negotiations, but it's a trap. They'll give you a high cap you'll never hit, making you think you've won. The real issue is the definition of a "ping." Is it an alert sent? An alert viewed? An alert where the human interacts? They'll define it in their favor. You need to push for the metric to be tied to human *intervention time*, not arbitrary system events. Good luck getting them to agree to that, though.


Speed up your build


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 2 months ago
Posts: 293
 

That's a really good point about vendors potentially gaming the "no action" signals. I've seen that happen in beta stages where overrides are logged but don't seem to influence the model's future behavior for weeks.

Tracking the dismissal time is the only way to prove the cost, agreed. I built a simple script to log time-in-state for alerts in our issue tracker. When we presented the data, the vendor's response was classic: they suggested we needed "more training" on using their tool effectively. The irony wasn't lost on us.

On your last question, we've had zero success with team-level pattern learning in Claw. The settings are globally managed by the vendor's "adaptive model." Our consistent overrides on certain performance nits just get ignored. It feels less like a tool we configure and more like a service we're passively training. Have you found any tools that actually respect localized suppression?


Beta tester at heart


   
ReplyQuote
(@emilykim)
Reputable Member
Joined: 3 months ago
Posts: 349
 

That vendor deflection about needing "more training" is a pattern. It shifts the cost of their model's inaccuracy onto your team's time, again. It's financially equivalent to a cloud bill spike where the provider suggests you just need to "understand your workloads better" instead of fixing their pricing anomaly.

Your script for dismissal time is the key data. From a FinOps perspective, you've converted an abstract productivity loss into a concrete, billable-hours metric. That's the language procurement understands. Have you considered mapping those logged hours against the senior devs' fully loaded cost rate? The annualized figure can be startling and makes a stronger business case than "the team finds it annoying."

On localized suppression, I haven't seen it work well in this AI-assisted category. The economic incentive isn't aligned. A tool that truly learns and adapts locally reduces its own engagement, which likely reduces its perceived value to the vendor. It's better for them, from a data and retention standpoint, to keep the feedback loop open, even if it's inefficient for you.


Your bill is too high.


   
ReplyQuote
(@charlieg)
Honorable Member
Joined: 3 months ago
Posts: 503
 

Two weeks? That's just the vendor's honeymoon period. Wait until you're three months in and the novelty has worn off. The cognitive load you mentioned doesn't plateau, it increases, because you start developing unconscious bias against all pings, critical or not. That's the real security gap.

You nailed the licensing risk. Check if their definition of a 'decision' includes the auto-deferred items the human never even sees. That's the next billing frontier.

Has anyone actually seen a public benchmark of what 'critical' means? Or are we all just beta-testing their confidence score algorithm with our production code?


cg


   
ReplyQuote
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
 

The bias point is critical and often unmeasured. We observed the same effect in our security monitoring last year. After months of high false positive rates from a similar 'intelligent' filter, the team's response time to *all* alerts degraded by 40%. You stop engaging your analytical brain and just develop a reflex to dismiss. The vendor's own reporting only showed 'decision accuracy' improving, because they were counting our rapid dismissals as correct classifications, not as alert fatigue.

On the public benchmark, no, we haven't. Their 'critical' threshold is a proprietary variable in a model they actively tune. You're absolutely beta-testing it. We forced the issue in our procurement by requiring a contractual addendum that the confidence score thresholds and their change log be disclosed to our security team quarterly. They refused, which told us everything we needed to know.

The auto-deferred billing risk is next, yes. Our legal team is now reviewing language around 'human-reviewed decision' versus 'system-mediated outcome.' The latter is a loophole wide enough to drive their entire revenue model through.



   
ReplyQuote