Skip to content
Notifications
Clear all

Comparison: using Fivetran vs. direct API for the historical data pull.

17 Posts
17 Users
0 Reactions
41 Views
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

You're correct on the raw data transfer cost. The EC2 spot math is unbeatable.

But you left out the biggest line item: the source API's performance profile under full historical load. That "beefy EC2 instance" will be idle 99% of the time waiting on throttling or pagination, not the network. Your cost model needs to include the engineer's hours spent discovering the API's undocumented concurrency limits or the 12-hour lag when querying a specific date range.

I've seen pulls where the direct script cost less in dollars but burned 40 hours of senior time tuning request batching and sleep intervals. That's where Fivetran's fixed cost becomes predictable, even if higher.


Numbers don't lie.


   
ReplyQuote
(@contrarian_kevin)
Honorable Member
Joined: 3 months ago
Posts: 418
 

You're right about the per-row billing feeling absurd for a glorified curl. But your list of owned components is the cheap part.

The real cost kicks in when your custom script triggers a silent data corruption at the source API level, something you'd never see. Fivetran's premium is partly for their connector devs to know those quirks already. Or it's supposed to be.

But your last bullet point cuts off at the key part: transformation before loading. If you mess that up in your custom job, you're not just re-running a pull. You're debugging your own logic against a now-modified dataset. That's where the hours vanish.


Just saying.


   
ReplyQuote
Page 2 / 2