Another year, another vendor decides the grass is greener in the Identity Resolution pasture. Snowplow, the “we give you the rails, you build the train” data collection platform, just announced their own identity stitching module. I’ve been cynically refreshing my CRM (currently HubSpot, last year was Salesforce, the year before that was a custom Frankenstein) for the better part of a decade, so forgive me if my immediate reaction is a long, slow blink.
The premise is classic platform expansion: you’re already piping your first-party event stream into their schemas, so why not let them handle the messy business of figuring out if `user_id: 123` is the same human as `anonymous_id: abc` from the mobile app? They’re promising deterministic rules, probabilistic matching, and of course, the ability to output the resolved identities back to your warehouse and activation destinations.
My skepticism stems from the fundamental tension in their model. Snowplow’s entire value prop has been control and transparency—you own the raw data, you define the schema, you manage the pipeline’s guts. Identity resolution is a discipline drowning in subjective decisions and hidden decay.
* Where do they land on the inevitable conflict between a high-confidence but stale CRM ID and a fresh but anonymous cookie?
* How is match key management handled when your marketing team decides that “email domain” should suddenly be a strong signal for B2B, flooding your customer graph with false company-wide linkages?
* Crucially, what’s the feedback loop for when their stitching inevitably creates a duplicate monster that triggers the same lead seven times in a minute? Can you audit, override, and see the actual logic trail, or are you just trusting a new black box?
I’m morbidly curious to see how this integrates—or more likely, competes—with the existing CDP and CRM tools in a stack. If I’m already using Snowplow events to build audiences in Segment, and those audiences sync to Salesforce for the sales team, and to HubSpot for marketing, where does this new resolved identity table live? Does it become the single source of truth, or just another opinion to be reconciled?
The announcement blog post is, predictably, light on the gritty details of failure modes and operational overhead. I’ve documented enough broken data marriages during my annual CRM migrations to know that the devil isn’t in the resolution algorithm; it’s in the daily maintenance of the rules, the cleanup of the bad merges, and the process for untangling the mess when sales submits a support ticket because their champion stakeholder just got merged with a low-level prospect from a competitor.
So, who’s actually trialing this? I want to hear from teams running it in parallel with another resolver (like Segment’s or a homegrown warehouse model). What broke in the first week? What, if anything, genuinely improved in your activation metrics versus your old method?
That tension is exactly why this move smells like lock-in in a fancy bottle. You trade their beautiful, transparent rails for a black-box stitching service where the most important logic - the decay rules, the conflict hierarchy - becomes proprietary.
Once your identity graph lives in their module, you're not just piping events anymore. You're asking for permission to understand your own customers. Good luck explaining your match rates during an audit when you can't point to the exact deterministic rule that failed.
Trust but verify – and audit
That's a good point about audits. Even if their logic is "open source," running your own instance to inspect it probably defeats the purpose of using their managed module.
Has anyone seen a vendor in this space actually publish their conflict resolution hierarchy as part of the product spec, not just a blog post?
You hit the core tension. Their whole selling point is that you own the infrastructure and see every component. This module inserts a proprietary black box into the middle of that.
The real cost isn't the licensing fee, it's the cloud bill for the compute needed to run their matching logic at your event volume. Probabilistic matching across years of user history is a massively expensive graph problem. They'll abstract that away, and you'll get a line item for "Snowplow Identity Compute" that's impossible to optimize because you can't see the queries.
You traded your "Frankenstein" for a vendor whose cost levers you can't even access.
cost optimization, not cost cutting
Okay, wait, can you explain the "control and transparency" bit a bit more? I'm new to this side of things. Our team uses Snowplow, and I thought the whole point was that we *don't* get stuck with black-box vendor logic.
If this stitching module is proprietary, doesn't that kind of break the original promise? We're paying to avoid that exact problem with our other tools.
That's a really sharp question, and I think you've put your finger on the exact worry a lot of us have. You're right that the original promise was about avoiding black-box logic.
The key distinction they'll likely make is that the core pipeline - your raw data collection, validation, and storage - is still open and yours. This module is an optional add-on that sits on top. So the promise isn't *broken*, but it's definitely being *tested*. You can choose not to use the module and keep full control, but then you're building that complex stitching logic yourself.
It does create a weird hybrid model. You get transparency for the data in, and potentially a black box for the identity graph out. For some teams, that trade-off for faster implementation might be worth it. For others, it defeats the purpose.
~Harry
The hybrid model you describe is precisely where architectural debt accumulates. If the identity module becomes a required source of truth for downstream models, you've effectively created a critical path dependency on a system whose internal state you cannot debug.
You can own the raw data lake, but if your analytics and activation layers consume the processed identity graph, any degradation in that black box corrupts every dependent system. The cost isn't just the module fee, it's the eventual need to build a parallel, transparent identity system to validate the vendor's output, which defeats the purpose of buying it.
infrastructure is code