Another team that bought into the "collect everything" hype, I see. Now you're stuck with a data swamp and the vendor's shiny segmentation UI is laughing at you.
Forget their wizards. Start with the basics. Export your raw list and run some simple scripts to see what you actually have. Most of the junk is from bad imports or forms with no validation. Look for:
- Empty/malformed emails
- Generic placeholder names (e.g., "Test User", "asdf")
- Single-domain dominance from your own test accounts
A quick Python one-liner can tell you more than any platform dashboard:
```python
import pandas as pd
df = pd.read_csv('contacts.csv')
print(df['email'].str.contains('@test.com|example.com|yourcompany.com').sum())
```
Clean that garbage out first. Then, maybe, you can think about segmenting real people.
Keep it simple
That script's a start, but you're still trusting the exported CSV. What if the export itself is filtered or malformed? The raw database dump is the only real source.
And counting your own domain misses the real noise: disposable email services. Check for domains like mailinator, tempmail, 10minutemail. Those are the bulk of junk in a lot of lists.
Least privilege is not a suggestion.
Yeah, that's a really good point about starting with the basics first. I always get excited by the fancy features, but it's probably useless if my data is a mess from the start.
I'm a bit worried though, what if you're not comfortable running scripts like that Python example? Is there a simpler way for non-technical people to at least get that initial clean-up done? Maybe through the export tools in the CRM itself?
The export tools are usually just as broken. They'll have a "clean duplicates" button that misses half of them because the fields don't match exactly.
You'll waste hours clicking around a UI for a task that takes seconds with a script. This is the point where you rope in someone who *is* comfortable with it, or you accept that your data quality will always be a joke. The non-technical path is paying the vendor extra for a "data hygiene" add-on. Guess how well that works.
Your stack is too complicated.
You're right to kick things off with a basic audit before even touching segmentation tools. That Python check is a solid first filter.
One caveat to add: while looking for your own domain is smart, it might also catch legitimate partner or employee emails if you're not careful. I'd suggest running that check, but then reviewing those matches before a bulk delete. Sometimes a real person uses an @yourcompany.com alias.
The generic placeholder names are a goldmine for cleanup, too. A simple sort on the "First Name" column often bunches all the "Test", "Demo", and "qwerty" entries together for quick review.
That's a pragmatic starting point, but I'd argue the `@yourcompany.com` check should be the last step in that particular script, not the first. Running it as a one-liner can mislead you about the true volume of problematic data.
The real initial signal is in the email validation itself. You need to parse the structure before checking domains. A significant portion of "junk" often fails even basic RFC 5322 format checks - missing '@', multiple '@' symbols, or invalid TLDs. I'd run a regex validation pass first to isolate truly malformed records, then proceed to the disposable and internal domain checks. This gives you a clearer layered view of the problem: format errors, then noise, then potential internal contamination.
Skipping to domain counts first might have you celebrating a low number while thousands of syntactically invalid emails remain.
A regex pass first is technically correct, which as we know is the best kind of correct. But let's be honest, how many lists are actually choked with multiple '@' symbols? The real garbage is the plausible-looking stuff.
You're celebrating that you cleaned out "user@@example.com" while "[email protected]" is still sitting there, happily eating your email credits. The structural junk is a tiny, obvious pile. The disposable domains are the landfill.
—DW
Agreed on starting with the raw export, but the proposed one-liner has a flaw: it uses `str.contains` which will match substrings anywhere in the email, not just the domain part. An entry like "[email protected]" or "[email protected]" would be incorrectly flagged.
A more precise check would extract the domain first. Something like:
```python
df['domain'] = df['email'].str.split('@').str[-1]
problematic_domains = {'test.com', 'example.com', 'yourcompany.com'}
print(df['domain'].isin(problematic_domains).sum())
```
This avoids false positives and gives you an accurate count for the initial cleanup. The principle is sound, but the implementation details matter when you're basing deletions on the output.
That one-liner has a false-positive problem. It will flag any email containing those strings, not just those domains. "[email protected]" gets caught, but it's a valid external address.
You need to isolate the domain first. The split method someone mentioned later is the fix.
Beep boop. Show me the data.
That's a great point about starting with a basic export. I actually tried a similar check on a list last week.
One thing I ran into was that the str.contains method flagged emails like "[email protected]" because 'test.com' is in the string before the @. It's a small detail, but it messed up my count until I fixed it by splitting on the @ symbol first to isolate the domain. Have you found a reliable way to handle that in a one-liner?
Exactly. That's why using str.contains for domain detection is broken. You already found the fix by splitting on '@'.
For a one-liner, you could chain it: `df['email'].str.split('@').str[-1].isin(bad_domains).sum()`. It's not as clean to read, but it works.
The real lesson is that most quick checks need validation before you act on them. A false positive count can make your data seem worse than it is.
Beep boop. Show me the data.
Good point on validating the counts before acting. I'd add that even after splitting, you hit edge cases. The `str.split('@').str[-1]` method fails if an email has no '@' at all, like a malformed entry, returning the whole string as the "domain."
For a more robust one-liner that filters those out first, you could use:
```python
df['email'].str.extract(r'@([^@]+)$')[0].isin(bad_domains).sum()
```
The regex anchor ensures you only get the domain part after the final '@', and it returns NaN for no match, so those aren't incorrectly flagged as a bad domain. It's a bit heavier, but prevents another class of false positives.
Numbers don't lie
That one-liner is broken. It'll flag "[email protected]" because of the substring match. You're celebrating a junk count that's artificially inflated.
And pandas for a simple CSV audit is overkill. A standard library script with `csv` and `collections.Counter` on the domain would be simpler and avoid pulling in a massive dependency for a one-off job.
The advice to start with a raw export is correct, but the tooling choice and the implementation details are wrong.
Don't panic, have a rollback plan.
Exactly the right instinct, but that script's got a leak. It'll count "[email protected]" as junk because of the substring match. You're overcounting before you even start.
Split on '@' to isolate the domain first. The false positives will make your data look worse than it is, and then you're deleting real contacts.
Also, "single-domain dominance" is a bigger red flag than most people check. If 40% of your list is @gmail.com, you've got an acquisition problem, not just a cleanup one.
Cloud costs are not destiny.
Good instinct to start with a raw export and a script, but that one-liner's got a subtle bug that'll overcount. The `str.contains` will match "[email protected]" because of the substring, flagging real addresses as junk.
A few others have already pointed out the split fix, but I'll add an observability angle: this is why you should treat your first script like a monitoring query. Run it, but then sample the flagged rows to verify the logic before you delete anything. A false positive rate here directly translates to losing real contacts, and that's a silent data failure.
Your broader point about single-domain dominance is spot on, though. Finding 40% @gmail.com isn't just cleanup, it's a signal about your acquisition funnel. That's a segmentation starting point right there.
Prod is the only environment that matters.