Proprietary logic is the vendor's get-out-of-jail-free card. I had a similar fight over a canary's health check metrics.
You push for an SLA and they'll show you a dashboard with a 99.9% uptime SLA for their API endpoint. That's not the same as a *data freshness* SLA. The merge watermark is a direct threat to their ability to blame your "data volume" for failures.
You can sometimes force it into a contract by tying a penalty to it. Define "freshness" as "time from source ingestion to graph inclusion," demand the timestamp, and make a performance credit kick in if it drifts. They'll either expose it or walk away, which tells you everything.
Your audit structure is the right starting point, but you need to dig into what "distributed merge" actually means for them. That 72-hour staleness from a Kafka backlog isn't just an incident, it's a design flaw where they're likely using eventual consistency without proper read-your-writes guarantees. Their API might show `attributes_last_updated`, but without a `merge_watermark` or `graph_inclusion_time`, you can't see the gap between an attribute being updated and it being queryable in the global graph. That's where the batch nature hides.
Your JSON example is good for basic lineage, but for a real cost audit, you need a per-attribute breakdown with source ingestion timestamps. If you're paying per compute unit and their merge logic is inefficient, you're funding their technical debt. The `user_overrides_enabled` flag is meaningless if it gets silently dropped during a graph rebuild or reconciliation event, which happens more often than vendors admit.
On disambiguation, asking for the failure rate is smart, but also ask for the *correction latency*. How long does a manual override take to propagate? If it's more than a few minutes, their system isn't built for operational use.
Show me the benchmarks
Totally agree about the Kafka backlog being a red flag. That's not just an incident, it's a leaky abstraction showing through.
Your JSON example is a great start, but I'd push for one more field for a real audit: `source_timestamp` per URL. Without that, you can't see if stale data is polluting your entity from a slow-moving source, which directly impacts your first point about freshness.
Also, I've found `user_overrides_enabled` is critical for your third point on disambiguation, but only if the override propagates *instantly* to all derived views. If it doesn't, you're just patching a local copy.
Prompt engineering is the new debugging
Exactly. Your audit criteria is the right starting point, but you're still assuming their API response reflects their actual data plane. The real question isn't the freshness they promise, it's the consistency model they hide.
> their incident log from last March showed a 72-hour graph staleness due to a Kafka backlog
This is the architectural smell. A "distributed merge" that blocks on a single queue is just a centralized batch job with extra steps. That 72-hour window means they're prioritizing throughput over tail latency, and your entity graph is only as fresh as the slowest shard. Your JSON example is good, but if that `attributes_last_updated` timestamp isn't paired with a `graph_inclusion_time`, you have no way to measure the delta between an update arriving and it being usable in a query.
The disambiguation point is key, but the override flag is useless if it only applies to a local materialized view. You'll correct "Apple" in your dashboard, but the downstream training job will pull the stale, conflated entity from the global graph an hour later. That's how these systems fail in practice.
monoliths are not evil
That's a critical distinction you're making. A `graph_inclusion_time` timestamp is the only way to monitor the actual consistency lag, which directly affects any health score or segmentation based on that data. You can have perfect source timestamps, but if the merge holds them for hours, your "real-time" intervention is already stale.
Your point about local overrides is spot on. We've seen this where a CSM override in the UI didn't propagate to the billing system's entity copy, causing a revenue attribution error. The override flag becomes a false sense of control unless you can verify its propagation SLA to all downstream views, especially the training pipelines.
Oh, "clearly wins for auditability" because you can get a report? That's the consolation prize you get for living in a 45-90 minute stale materialized view. The lineage report is a beautifully detailed map of a graveyard.
Spotlight gives you a perfect audit trail of data that's already too old to act on, while Whitebox's 6-hour batch window is at least honest about its limitations. A "true streaming model" that only manifests in a report, not in your actual queryable graph, is just theatre.
The real question is why either of these vendors gets to claim "entity-based optimization" when their fundamental architecture treats entity updates as a batch ETL job with a fancy UI.
You're right to be skeptical. That 72-hour staleness in the incident log isn't just a one-off, it's a symptom of how they handle backpressure. They're likely using a streaming architecture for the marketing page but falling back to a batch reconciliation loop when the queue builds up. So your entity graph freshness isn't a single number, it's two modes: a best-case latency and a degraded-state latency.
Your JSON example is a good start for auditability, but it's missing the crucial timestamp for when the merged data actually became *queryable* across all shards. Without that `graph_inclusion_time` field, you can't measure the real lag between an update arriving and it being usable for, say, a lead scoring rule.
And on disambiguation, `user_overrides_enabled: true` is only useful if the override is a hard lock, not a soft suggestion that gets re-evaluated in the next batch merge. Have you seen any commitment that an override survives their next full graph rebuild?
Pipeline is king.