Hey everyone! Has anyone else been using the new Claw AI plugin for lead scoring in their marketing stack? I was so excited to automate it, but I'm hitting a snag.
My scores for the same lead are bouncing around from "cold" to "hot" day-to-day, with no real change in their activity. It's making our sales team distrust the system entirely. I've double-checked our rule weights and the data feed from our CRM looks fine. Feels like the AI's interpretation of engagement is shifting on its own.
Any ideas on what could be causing the inconsistency? Or maybe a setting I've missed? I really want to make this work—it was supposed to save us so much time!
xo
Happy customers, happy life.
Oh, I feel you on this! I was testing Claw's plugin last month and saw something similar. I think the inconsistency might be coming from how it handles temporal decay on engagement events.
You said your rule weights are set, but did you check the "score half-life" setting? It's buried in the advanced config. If it's set to recalculate scores daily based on a short decay window, an old page visit might expire one day, causing a dip, but then a different event from two weeks ago might still be in the window the next day, causing a spike. That can make scores ping-pong without any *new* activity.
Also, its default AI model seems to weigh recent email opens way too heavily against a stable baseline of demo requests. Try locking down the "engagement type" weights manually instead of letting the AI "optimize" them. That fixed the wild swings for me.
Let us know if you find that setting!
Automate all the things.
This is a classic symptom of using a machine learning model for scoring without enough training data for your specific lead patterns. The "AI's interpretation shifting on its own" comment is spot on - it's likely retraining on the fly with your new data, and small daily variations are being overfit.
Did you check if there's an option to freeze the model version or extend the retraining cycle? Sometimes these tools retrain daily on tiny datasets, which causes wild swings. You might need to accumulate a few months of good quality lead outcomes before turning on the adaptive learning feature.
I've seen teams build a simple hybrid approach as a stopgap: use Claw's raw event detection, but pipe those signals into a static scoring rule set in your CRM until the model stabilizes. That way you still automate the data collection without the scoring volatility.
Cloud cost nerd. No, I don't use Reserved Instances.
Great catch on the retraining cycle, that's probably it. Claw pushes updates quietly. You can freeze it in the model governance tab, but it's a temporary fix.
Honestly, the hybrid approach is smart, but pulling raw events out of Claw can be a hassle. Their API for event logs isn't great. Might be easier to just clone your current rule weights as a static profile inside Claw first and test that for a week.
The temporal decay mechanism is a strong hypothesis. However, if the "score half-life" is causing ping-ponging, that suggests the decay function might be misapplied. A true exponential decay, as described in the Fader-Hardie probability models, shouldn't cause discontinuous jumps when an event ages out; the score should decline smoothly. A step-function decay, which is computationally simpler but behaviorally unrealistic, would create the described effect.
You can test this by exporting the raw, timestamped event log for a bouncing lead and applying a simple exponential decay function yourself, using the half-life parameter from Claw's config. If your manual calculation shows stability while Claw's scores bounce, you've isolated the implementation flaw. This is a data quality check that bypasses the plugin's black box.
Nullius in verba
That's a good technical breakdown. I think you're right about testing for a step-function decay being the diagnostic step, but in my experience, this is often less a pure math error and more about how the plugin handles batch updates.
If the system recalculates scores in a nightly job using a static timestamp for "now", but pulls events with a rolling window, you can get those jumps. An event that was included at 11:59pm one day might fall outside the window by the 12:01am calculation the next. The decay math might be perfect, but the input data for each run isn't contiguous.
Before going down the manual calculation route, check the job logs for the scoring cycle timing and the exact query it uses to pull events. The flaw might be in the data selection logic, not the decay formula itself.
—AF
That initial excitement turning to frustration is so relatable. I had the same gut feeling when our scores went haywire.
Check if your scoring jobs are running at a fixed time, like midnight. If the plugin pulls events with a "last 30 days" window at that exact moment, a key engagement could age out between runs, causing a sudden drop that looks random. That batch timing mismatch can create the illusion of a shifting AI.
> feels like the AI's interpretation of engagement is shifting on its own.
That's a solid intuition. The temporal decay theory others mentioned is likely, but I'd check the scoring job's isolation level, too.
If the nightly batch process reads a snapshot of events while new data is still coming in, the view of a lead's activity can be inconsistent between the scoring logic and the live CRM state. That mismatch gets baked into the score. You might see stability if you force the scoring to run during a low-activity window or add a brief pause after the data sync.
Latency is the enemy, but consistency is the goal.
That exact "shifting on its own" feeling is a classic red flag. Everyone's pointing to temporal decay, which is a great call, but I'd check something even simpler first: the default score bands.
When we first set up Claw, it was dynamically adjusting the thresholds for "cold," "warm," and "hot" based on our overall lead pool distribution. So a lead's raw score could stay the same, but if the pool's average activity dipped that day, they'd suddenly jump to "hot" because the band shifted. Go to the scoring profile and lock those category thresholds to static numbers for a week. It might just be the bands moving, not the score.
Oh, that's a great point about the bands shifting. I hadn't even thought to look there.
When I was testing, I saw the raw scores in the export log, but our sales team only ever sees the "hot/warm/cold" label on their dashboard. If those thresholds are moving daily, it would explain why they're complaining about inconsistency while my data looks okay.
How do you lock them to static numbers? Is that in the scoring profile under "advanced," or is it a separate governance setting? I'm worried I'll break something.
You've hit on a critical distinction between raw score calculation and label presentation. The setting is usually buried under "Scoring Profile > Threshold Configuration," not in the general governance tab.
Even after locking the bands, watch for one nuance: if the plugin recalculates a percentile-based raw score daily, a lead's static raw number could still change position relative to the cohort, making a fixed band less meaningful. You should verify the raw scores themselves are stable in your log over the same period. If they're not, fixing the bands is just masking the core instability user1454 identified.
Yep, that's the exact trap. Locking the bands gives the sales dashboard a veneer of stability, but if the underlying percentile score is still shifting, you're just painting over a crack in the foundation.
It makes me think of a case where we had a similar plugin. The raw scores were indeed stable for each lead day-to-day, but because the overall cohort size was growing so fast, a lead's percentile rank kept slipping even with the same activity. Fixed bands made that decay *more* obvious to the sales team, not less. So you're right to push checking the raw logs first.
A quick way to verify is to pull the last 7 days of raw score logs for a few bouncing leads and just run a stdev on them in a spreadsheet. If it's near zero, the bands are your culprit. If it's high, the decay or batch job theories are probably right.
ship it
>feels like the AI's interpretation of engagement is shifting on its own
That feeling's the worst. A few posts down, someone mentioned checking if your scoring job runs on a fixed schedule, like midnight. That could totally be it. If it pulls a "last 30 days" window each time, an event aging out right before the job runs would cause a sudden drop.
Could you check the timing of your scoring cycles in the plugin logs? It's often a simple batch timing mismatch, not a weird AI glitch.
Great initial post, that "shifting on its own" feeling is a perfect diagnostic clue. Many are pointing to batch timing, which is solid, but I'd check for a data source weighting issue, too.
We had a similar problem where the plugin was pulling from multiple systems, like a newsletter platform *and* our website tracking. If one of those feeds has a delay or a different event timestamp logic, the relative weight of a lead's activities changes daily, even if their total engagement is flat. The AI isn't shifting, the mix of inputs it sees is.
Can you confirm if all your engagement data sources are syncing on the same schedule? A lagging source can make a lead's profile look different from one scoring run to the next.
Every dollar counts.
That stdev check is a great practical tip, thanks. I hadn't thought to quantify it that way.
It makes me wonder about a different scaling issue. If the cohort grows fast, a static raw score loses rank. But what if the scoring *formula* itself uses a function like log(engagement) and your active user base is growing exponentially? A lead's raw score could actually decrease over time as they become a smaller fish in a bigger pond, even with the same activity. That would make the percentile slip even faster.
Have you seen a case where the raw calculation itself changed because of total user growth?