You've put your finger on the crucial limitation. That's exactly why I layer in a statistical distribution check after the basic CLI validation. The tool's diff is necessary but insufficient.
We ran a similar migration and the CLI gave us a clean report. Only when we stratified the validation by user cohort and ran a Kolmogorov-Smirnov test on event property distributions did we spot the skew. The destination sampling was discarding events from a specific geographic region's mobile app at a slightly higher rate due to a sessionization nuance. Total volumes matched, but the composition was wrong.
Our false-negative rate for that specific class of "invisible skew" was around 18% using just the row-by-row diff. The CLI caught the obvious gaps, but the silent distortions required a different statistical lens.
Absolutely love the "map your five most complex event types first" advice. That's the exact kind of pragmatic, scars-earned wisdom that saves a project. We used a similar rule of thumb, but we called them "tracer events." You pick a few that travel the longest path through your old system, with the most transformations, and validate those end-to-end.
My one caveat to your approach: sometimes that low-volume, bespoke event is the "Refund" or "Chargeback" type. It's rare, but if it's broken, the financial impact is massive. We learned to categorize by business risk, not just volume. If reconiling it is a three-day rabbit hole, you're right to let it go. But if it touches money or a core compliance requirement, you might have to eat the time cost.
How do you decide which edge cases are truly "low-value"? Was it purely a volume metric for you, or did other factors creep in?
Implementation is 80% process, 20% tool.