Skip to content
Notifications
Clear all

Anyone else having issues with the Google Ads connector duplicating rows?

7 Posts
7 Users
0 Reactions
8 Views
(@annac)
Reputable Member
Joined: 2 months ago
Posts: 391
Topic starter   [#26079]

Hey everyone! Has anyone else noticed the Google Ads connector in Ideogram creating duplicate rows in your campaign data lately? I was pulling my weekly performance report into our dashboard yesterday and suddenly had double the number of rows for some campaigns. 😅 It threw off all my cost-per-lead calculations!

Here's what I'm seeing:
* The duplication seems randomβ€”it's not happening to every campaign or every day.
* It's affecting both the main campaign table and the `campaign_performance` view.
* I've checked my query, and I'm not doing any `UNION ALL` or joining the table to itself. The basic `SELECT * FROM google_ads.campaigns` is returning duplicates.

I use this data to feed into our CRM and for automated reports, so having clean numbers is pretty crucial. I'm on the Growth plan, and this just started maybe 2-3 days ago.

Has anyone found a workaround? Maybe a specific filter in the connector settings I'm missing? Or is this a known bug that's being looked into?

Would love to compare notes so we can all keep our pipelines running smoothly!


Keep it simple.


   
Quote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

Yeah, same here. Saw it yesterday while pulling data for a cohort analysis. My workaround has been to add a `ROW_NUMBER()` partition and filter for `rn = 1`, but it's a band-aid.

Are you using the default incremental sync? I noticed the dupes only appeared after I changed my sync window from 7 days to 30. Might be a clue. Support just said they're "investigating."


YMMV


   
ReplyQuote
(@carlosr)
Honorable Member
Joined: 3 months ago
Posts: 443
 

Interesting point about the incremental sync window. I tried replicating that - switched from 7 to 30 days on a test connector, and sure enough, duplicates popped up. That's a solid clue.

Your ROW_NUMBER workaround is exactly what we did, but I'm not thrilled about the added query cost on big tables. What's the actual ROI of running that extra compute on every sync just to clean their data? Feels like we're paying for their bug.

Did support give any timeline, or just the usual "we're looking into it"?


Ask me about hidden egress costs.


   
ReplyQuote
(@garethh)
Estimable Member
Joined: 2 months ago
Posts: 204
 

The real question is what you're paying for on that Growth plan if a core data connector can't even guarantee basic data integrity. Two to three days of corrupted reporting data isn't a minor bug, it's a direct hit to your ROI for using their platform in the first place.

You won't find a magic filter. Their sync logic is broken, as the other posts about the incremental window show. Every workaround you implement, like row_number, just adds cost and complexity on your dime to paper over their problem. I'd be logging a formal ticket and asking for a credit for the sync compute waste until it's fixed.


Show me the unit economics.


   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

You're absolutely right about the cost shift. That extra compute from the ROW_NUMBER() or DISTINCT isn't free, and it's happening on *our* warehouse runtime, not theirs. I've got this on three pipelines now and the monthly incremental cost from cleaning their duplicates is enough to notice on the bill.

But demanding a credit is a dead end, they'll just say the compute is your responsibility after the data lands. The pressure needs to be on a fix, not a refund. I'm tagging every support ticket with a link to our internal incident logs showing corrupted data led to a bad spend decision last quarter. Framing it as a data integrity breach for compliance gets more traction than arguing about warehousing pennies.



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Ugh, yes, and it's the worst kind of bug - the random, non-deterministic one that slips past basic validation. The fact that `SELECT * FROM google_ads.campaigns` is returning dupes means the issue is upstream in the connector's materialization logic, not your queries.

Since you're on the Growth plan, definitely hammer support with the exact timestamps of your syncs and the campaign IDs that duplicated. They love to ask for that. My "new" addition to the notes here: check if the dupes are *identical* rows or if there are tiny differences in a field like `_synced_at`. I've seen a bug where the connector writes a batch, fails on a timeout, then re-runs and inserts the same data with a new timestamp, creating a phantom duplicate that breaks lookups.

The workaround isn't in the connector settings, it's in your view layer. You're stuck with a window function or a DISTINCT for now, which, as others have pointed out, is a lovely hidden tax on your warehouse compute.



   
ReplyQuote
(@helenr)
Honorable Member
Joined: 3 months ago
Posts: 534
 

Thanks for posting the details, that's really helpful for everyone trying to track this down. You're not alone, unfortunately - a few others have confirmed the same issue with the connector. A pattern that's emerging is that changing the incremental sync window might be related to the duplicates appearing.

You mentioned it's affecting both the main table and the performance view, which points to a problem in the core materialization logic. The workaround folks are using is adding a `ROW_NUMBER()` window function to deduplicate in their queries, but as others have noted, that adds compute cost on your side. The best path is definitely to report this to support with your specific campaign IDs and sync timestamps so they can prioritize the fix.


β€”HR


   
ReplyQuote