I like this approach, but I'm worried about that regex. What happens if the filename has a hyphen or a number? I'm still learning but I thought regex could be tricky.
That's a great worry to have, because it's exactly what happens! You're right on the money about regex being tricky. If your regex only looks for letters `[a-z]`, a filename like `data-migration-2025.pdf` will just vanish from your count, because the hyphen and numbers don't match the pattern.
It's one of those things that seems to work perfectly... until it doesn't. I learned this the hard way during a migration project. We were tracking PDFs for a campaign, and halfway through, marketing updated all the filenames to include the quarter. Suddenly our download numbers looked amazing - because we were only catching the old files with the simple names, and none of the new ones!
If you're stuck using regex, you can make it a bit safer by expanding the pattern to allow for more common characters. Something like `^/whitepapers/[a-zA-Z0-9_.-]+.pdf$` would catch hyphens, numbers, underscores, and dots in the name part. But honestly, it's a band-aid. As soon as someone uses a parenthesis or an emoji (yes, I've seen it!), it breaks again.
Backup first.
The sentiment is right, but that regex is a time bomb. It assumes a clean path and will choke on query strings, URL encoding, or even a simple uppercase 'PDF'. You'll get NULLs for whitepaper_name and never know until a report looks off.
Also, parsing daily is fine until sales needs to see campaign performance with a shorter lead time. Then you're stuck explaining why their dashboard is stale. Server logs are the foundation, but calling third-party tracking "drama" ignores the real-time validation it forces you to build.
Your CRM is lying to you.