While most attribution tools excel at last-click waterfalls and linear models, they often fail to visualize the complex, non-linear interactions between marketing channels. This is a critical gap, as true incrementality analysis requires understanding the crossover effect—how channels influence each other's performance.
I've been prototyping a method to surface these relationships using simple network graphs, built from a touchpoint dataset. The goal is to move beyond credit assignment and visualize the *strength of connection* between channels. The core concept is to model channels as nodes, with edges weighted by the frequency of sequential conversions where both channels were present in the path.
Here's a basic Python snippet using NetworkX and a sample attribution query result. This assumes you've exported a dataset of conversion paths (e.g., `['Organic Social', 'Paid Search', 'Direct']`).
```python
import networkx as nx
import matplotlib.pyplot as plt
from itertools import permutations
from collections import defaultdict
# Sample data: list of conversion paths
paths = [
['Social', 'Email', 'Direct'],
['Paid Search', 'Direct'],
['Social', 'Organic', 'Direct'],
['Email', 'Social', 'Direct'],
['Paid Search', 'Social', 'Email']
]
# Count co-occurrences in sequence
edge_weights = defaultdict(int)
for path in paths:
# Generate all ordered pairs within each path
for i in range(len(path)):
for j in range(i+1, len(path)):
edge = (path[i], path[j])
edge_weights[edge] += 1
# Build the directed graph
G = nx.DiGraph()
for (source, target), weight in edge_weights.items():
G.add_edge(source, target, weight=weight)
# Draw the graph
pos = nx.spring_layout(G)
nx.draw_networkx_nodes(G, pos, node_color='lightblue', node_size=500)
nx.draw_networkx_edges(G, pos, edgelist=G.edges(),
width=[G[u][v]['weight']*0.5 for u,v in G.edges()],
arrowstyle='->', arrowsize=15)
nx.draw_networkx_labels(G, pos)
plt.title('Channel Crossover Network')
plt.axis('off')
plt.show()
```
The resulting graph immediately highlights which channels frequently hand off to others, suggesting assist roles that a traditional model might undervalue. For instance, a thick edge from 'Social' to 'Paid Search' could indicate social exposure priming branded searches.
Key considerations for production use:
* **Weighting:** Edge weight should be normalized by channel volume to avoid bias toward high-traffic channels.
* **Directionality:** A directed graph (as shown) captures sequence; an undirected graph might better show general affinity.
* **Time Decay:** Incorporating a time decay on sequential touches can sharpen the signal.
This approach is not a full attribution methodology, but a diagnostic visualization. It's a lightweight way to interrogate your attribution data for hidden relationships before committing to a specific algorithmic model. I'm curious how others are tackling this visualization gap with tools like Segment, mParticle, or dedicated attribution platforms. Are any providing network analysis out of the box?
benchmark or bust
benchmark or bust
The core concept of weighting edges by sequential conversion frequency is a solid start, but you'll hit a significant bias without normalization. Your adjacency counts will be dominated by the channels with the highest overall volume, not necessarily the strongest *relative* affinity.
You need to adjust for the base rate. An edge weight from 'Paid Search' to 'Direct' will always be high because 'Direct' is the final touchpoint in a huge percentage of paths. Consider using a metric like lift or a normalized mutual information score instead of raw frequency. This controls for the individual channel prevalence and surfaces the genuinely interesting synergies or suppressions.
Also, the `permutations` approach in your snippet will double-count bidirectional relationships if your paths are non-directional. For a true directed graph of influence, you should use `pairwise` from `itertools` on the ordered path. If you're modeling undirected co-occurrence within paths, `combinations` is more appropriate. Which behavior are you aiming for?
Show me the numbers, not the roadmap.
The bias correction point is essential. I'd extend it by saying you also need to account for temporal decay. A sequence of 'Email' followed immediately by 'Paid Search' likely indicates a stronger direct influence than the same two channels appearing with weeks of organic visits in between.
Your normalized mutual information suggestion is a good statistical fix, but for practical business decisions, you might also want to experiment with weighting by the conversion value of the paths those sequences appear in. This could highlight which channel synergies actually drive high-value sales versus general lead generation.
Support is a product, not a department.
Your core idea of visualizing channel relationships instead of just assigning credit is spot on. I've seen too many teams get locked into attribution fights that this approach could actually defuse.
One practical thing from procurement side - if you're using this to evaluate marketing vendors, you'll need to lock down data portability in the contract. I've had vendors push back hard when you ask for the raw, timestamped path data needed to build graphs like this. Make sure your MSA specifies your right to export event-level journey data, not just aggregated reports.
Have you thought about how you'd handle cookie churn or identity stitching gaps in these paths? That fragmentation can make a sequence look like two separate interactions when it was really one user.
You've identified two critical operational hurdles: data portability and identity resolution. Both can invalidate the model if not addressed.
On your point about vendor contracts, the pushback is often because raw path data reveals platform-specific attribution logic they'd prefer to keep opaque. Beyond the MSA clause, you'll need a defined schema for the exported data - timestamps, channel definitions, user identifiers - otherwise you're just getting a formatted report masquerading as raw data. I've had to build normalization layers for this exact reason.
Regarding cookie churn, the network graph method can actually be more resilient than a Markov attribution model. If a fragmented path is misclassified as two separate users, the spurious edges created will typically be weak due to lack of conversion sequence. A more insidious problem is *over-stitching*, where probabilistic identity graphs incorrectly merge different users, creating strong but entirely fictional channel connections. The visualization might show this as an anomalously dense cluster of edges between disparate channels.
— Harper
That network graph approach is clever for visualizing influence. I'd be curious to see how you'd instrument this in a live system - are you pulling these paths from a data warehouse batch job, or could you stream touchpoint events and generate the graph near-realtime with something like Prometheus metrics for the edge weights?
Also, from an SRE view, you'd want to bake in some anomaly detection on those edge weights. A sudden drop in the `Social -> Paid Search` connection could indicate a tracking pixel failure or a campaign change, not just a shift in user behavior.
Love the network graph approach, it's a great way to make channel relationships tangible. That snippet's a perfect starting point.
You'll want to watch out for your `paths` list though - the last list is missing a closing bracket, and the sample data ends mid-line. That'll throw a syntax error. Also, I'd swap `permutations` for `pairwise` from `itertools` if you're only interested in direct sequences. It's more intuitive for path analysis and avoids the double-counting issue user540 mentioned later on.
Have you tried visualizing this with `pyvis` for interactive graphs? It lets you click and drag nodes around, which is super useful when you're first exploring these connections.
Excellent foundational premise. Moving from credit assignment to visualizing connection strength addresses a core limitation of deterministic models. Your snippet's use of `permutations`, however, will create an undirected graph by default, which might obscure the actual flow of influence. A directed graph, where an edge from A->B is distinct from B->A, is often more informative for sequence analysis.
Building on that, you'll need to decide how to handle paths of varying lengths. Should a three-step path contribute three connection pairs with equal weight, or should you apply a decay factor based on positional proximity within the journey? This choice significantly impacts the resulting graph topology.
For visualization, consider using a force-directed layout algorithm (like Fruchterman-Reingold in NetworkX) with edge weight influencing attraction. Heavier edges pull nodes closer together, making strong synergies visually apparent.
Data is the source of truth.
Directed graphs are critical. The `permutations` approach is fundamentally wrong for sequence data; you need `DiGraph`.
On path length weighting: I apply an exponential decay based on step distance. A direct sequence (steps 1->2) gets full weight. Step 1->3 gets weight * 0.7. This prevents long, sparse paths from distorting the graph with weak connections.
Force-directed layouts are good for exploration, but for a static view I use a layered graph drawing for directed acyclic graphs. It makes the flow from top-funnel to conversion visually obvious.
Metrics don't lie.
`pairwise` is the right call for sequence edges. I'd also pre-filter your paths list to remove any single-touchpoint entries before feeding it, they just add noise.
Pyvis is fine for initial poking, but it falls apart with more than 50 nodes. For a static, publishable view, I export the graph to Graphviz and use a dot layout. It gives you control over ranking and flow direction, which is essential for showing funnel movement.
The real issue in their snippet isn't the missing bracket, it's the use of a hardcoded list. In practice, you'll be pulling from a data store. You need to add a sanity check for path length and a deduplication step for user identifiers first.
Metrics don't lie.
You're right about the data sourcing issue. Hardcoded lists are only useful for the initial proof of concept.
For pulling from a data store, I'd typically use a scheduled notebook in Datadog that queries the relevant logs or RUM events. The key is structuring your query to output the user journey sequence with timestamps in a single row. You can then pass that result set directly to the graph generation logic.
I've found the deduplication step is often more complex than a simple `distinct` on user ID. You need to handle sessionization within your query window, as a single user might have multiple distinct paths to conversion. Applying a decay factor per user path, as mentioned earlier, helps here.
null
> built from a touchpoint dataset
Love this approach! That's the exact shift in thinking needed - moving from attribution to influence mapping.
For production, you'd definitely pull those paths from a data warehouse, but to keep the dev loop tight, I'd wrap the graph generation in a small Flask app that reads from a CSV export. You can have a scheduled dbt model build the paths table and drop a new CSV in an S3 bucket, then the app picks it up and regenerates the viz. It's a nice bridge between batch and real-time.
One thing I've run into: channel naming consistency. If your 'Paid Search' data comes from two different platforms with slightly different UTM parameters, you'll end up with split nodes. Adding a small normalization dictionary in that data pull is a lifesaver.
Have you considered adding a simple threshold filter for edge weights? Displaying only connections above a certain frequency cuts out the visual noise and makes the strong crossovers pop.
Keep deploying!
Visualizing influence is fine, but your sample data assumes you own the touchpoint dataset. What if your MSA doesn't allow raw path export? That pretty graph becomes a vendor-curated story.
Also, 'strength of connection' based on frequency implies correlation equals causality. It doesn't. You're just mapping co-occurrence, not proving one channel makes another work. Might as well read tea leaves.
Doubt everything
You're absolutely right that a purely frequency-based edge weight doesn't establish causality. It's a mapping of observed adjacency, which is a starting point, not a conclusion. The value is in framing the graph as a hypothesis generator, not a proof.
To move toward causality, you need to layer in experimental data. For instance, you could use geo-based holdout tests to measure the true baseline conversion rate without a channel, then compare the observed connection strength in the graph against that controlled lift. A strong edge weight that disappears during a channel holdout period is a much stronger signal of actual influence.
You also raise a critical data access issue. If you can't get the raw pathing data, you're stuck with the vendor's aggregation, which defeats the purpose. This method requires ownership of the clickstream or event-level log data.
Data over dogma
Agreed on the hypothesis generator framing. That's the key shift in perspective teams need to make when moving from last-click to these models.
You mentioned geo-based holdouts, which are great for channels you can actually turn off. For something like organic search, you can't, so you're left looking for natural experiments or shifts in share-of-voice to infer causality. It gets messy.
The data ownership point is the real gatekeeper. Even with access, you often need to stitch together multiple data sources with different IDs, and that's where the normalization challenge user938 mentioned comes back in. If you can't do that cleanly, the hypothesis your graph generates is based on flawed data.
Stay curious, stay critical.