Deterministic matching uses exact, known identifiers. Think email, user ID, phone number from a login. Two records match only if the identifier values are identical. It's precise but limited to first-party data you control.
Probabilistic matching uses statistical models on a set of attributes. Browser fingerprint, IP, device type, behavioral patterns. It calculates a likelihood of a match. Expands reach but introduces uncertainty.
Key operational differences:
* **Accuracy vs. Scale**: Deterministic is high-confidence, low-coverage. Probabilistic increases coverage at the cost of confidence intervals.
* **Data Requirements**: Deterministic needs a stable, shared key. Probabilistic needs a large volume of interaction data for the model to be meaningful.
* **SLO Impact**: Your identity resolution SLOs must differ. Deterministic can target 99.9% accuracy. Probabilistic requires a confidence threshold (e.g., 95%) and you must track false positives.
Most production CDP implementations use a hybrid. Deterministic for known users, probabilistic to stitch anonymous activity before a login. The weighting depends on your tolerance for error in audience activation.
—D
Five nines? Prove it.