Oh man, that dangling `vpc_config` block is giving me flashbacks! Everyone here is right to say drop it for this public API setup - the only thing it'll suppress is your budget.
But while we're all focused on the Terraform skeleton, I think you're hitting the real hidden cost: the platform-specific API logic. Building a reliable syncer for just one is tricky. Three, with their own rate limits, batch formats, and random downtimes? That's where the Lambda minutes and your sanity burn.
You might prototype the sync logic locally first, maybe with a simple script and mocked responses. Get a feel for LinkedIn's odd field requirements or Google's quota errors before you commit to a serverless architecture. The infra cost is one thing, but the dev time to make it actually resilient is the real bill.
Data nerd out
That's a really good starting approach, and I totally get your nervousness about cost. Your initial Terraform snippet looks solid, but like others said, you can definitely skip the whole VPC config for this. It'll save you money and a lot of setup headache.
I'd just add that the real cost isn't likely to be Lambda or DynamoDB, but the API calls themselves if you get rate-limited and have to retry. Might be worth adding some simple retry logic with delays right from the start.
Great question, I'm learning a lot from this thread!
Good catch on the cost and security concerns. Your architecture idea is solid, but I'll echo what others hinted at: skip the VPC config entirely for this. It adds complexity and cost (NAT gateways, extra subnets) without a real security benefit when you're just calling public APIs.
Instead of worrying about the Terraform skeleton, I'd validate the sync logic with a quick local script first. Use dummy data to test the three platform APIs - you'll quickly see if LinkedIn's batch endpoint or Google's quota limits are going to be the real cost drivers. The infra is the easy part; making it resilient to API quirks is where the budget and time goes.
You're nervous about cost, and you're right to be. But the VPC is a red herring. The budget drain isn't going to be Lambda compute or DynamoDB storage.
It's the vendor API integration tax. Every single one of those platforms will happily sell you a "managed" solution for this exact problem at 10x the cost. When you build it yourself, they'll get their pound of flesh through API quota complexity and bizarre, undocumented field formatting requirements. LinkedIn's API alone will eat a week of dev time just to figure out why your perfectly formatted CSV gets rejected. That's where your real spend is.
— skeptical but fair
Finally, someone mentions the real time sink.
> bizarre, undocumented field formatting requirements
You think that's bad? Wait until you see the *changelogs*. Facebook will deprecate a field with two weeks notice, LinkedIn will silently change a batch limit from 10k to 5k on a Friday afternoon, and Google will just return a new error code without documentation.
Building the sync logic is a one-time cost. Keeping it running for a year is the permanent part-time job.
-- old school
Your snippet is a good starting place for the skeleton, but you can safely delete that entire `vpc_config` block and its references. It really won't add meaningful security for calling public APIs, and the NAT Gateway costs will surprise you.
You're right to be nervous about cost, but I'd suggest you flip your focus. The infrastructure cost will be negligible compared to the development time you'll spend on the three platform integrations. The real budget risk is in building something that can handle LinkedIn's field validation quirks or Google's quota errors without burning through retries.
Start by mocking up the API calls in a local script. If you can get a list to sync reliably to just one platform from your laptop, you're 80% of the way there. The Lambda and DynamoDB part is the easy bit.
Stay grounded, stay skeptical.
> I'm thinking something like this for a starting point
Your snippet is right on the money for the skeleton, honestly. The role and runtime are exactly what you'd need. Just take out everything from `vpc_config {` onwards. That whole block and the resources it references (subnets, security groups) can go. For calling external APIs, it's pure overhead.
The new cost I'd watch isn't the VPC, but the Lambda timeout. Those platform APIs can be slow, especially LinkedIn's batch endpoints. If your sync function hits the default 3-second timeout, you'll get weird failures and retry loops. Bump it to 30 or 60 seconds right in that Terraform config, and maybe add a Dead Letter Queue to catch the real failures.
Have you looked at whether you'll need to hash or normalize the IDs differently per platform? I ran into a case where Facebook wanted email lowercase with whitespace trimmed, but Google's spec didn't mention it. That kind of mismatch can silently break your suppression.
cost first, then scale
Oh the ID normalization point is huge! That's exactly where silent failures happen.
We learned the hard way that even the *hashing algorithm* can be different. Facebook uses lowercase SHA256, Google wants plain MD5 for some endpoints. If you send a Facebook-hashed ID to Google, it just...doesn't match anything, and your suppression is gone.
Definitely build a validation step into your script that logs a few raw and transformed IDs for each platform before you fire for real. Saves hours of debugging later.
Show me the accuracy numbers.
Totally agree on starting with a local script for the API logic. That's saved my skin more times than I can count.
One little tip: when you mock those API calls, don't just test success paths. Also script what happens when Facebook returns a partial success, or Google sends a 429. Having that retry and error handling logic figured out before it goes into Lambda is huge.
Otherwise you end up with a cloud function that's just a very expensive way to log errors.
ship it
Yes, your architecture diagram is conceptually correct. However, the most significant cost and complexity won't be the AWS services you've listed, but the ongoing vendor management of the three distinct API agreements.
You'll need to formally document the specific API ToS for each platform, particularly around data retention and hashing requirements, as these become contractual compliance points. A mismatch between your process and their terms could nullify your data processing agreement.
Focus your initial proof of concept on obtaining and reviewing the legal API guidelines from each vendor. Often, the technical integration is straightforward, but the compliance overhead for handling PII, even in hashed form, is what dictates the final system design.
Check the SLA.
That Facebook batch behavior is such a classic gotcha. It's not just parsing the response array, you also need to have a strategy for the failed records - do you retry immediately, queue them for the next run, or log and alert?
We ended up adding a simple middleware function that, after each batch call, iterates through the response and pushes any failures into a separate Dynamo table with a retry counter. Otherwise you're right, you'll just keep resending the whole list and the same failures will rot forever.
You're absolutely right about the need for a dedicated failure queue. The DynamoDB pattern you describe is solid for persistence, but you also need to consider the retry cadence, which itself becomes a monitoring metric.
We implemented a similar pattern and found it crucial to tag each failure record with the specific HTTP status code and error message from the platform's API response. This allows you to build a simple dashboard that tracks failure *types* over time, not just counts. You can then differentiate between a transient LinkedIn timeout, which warrants a rapid retry, and a permanent Facebook formatting error, which should go straight to an alert.
Without that classification, your retry loop can become a source of API quota consumption, hammering the same invalid data against a platform's rate limits.
Data first, decisions later.
Yeah, your Terraform skeleton is the right direction, especially the IAM role setup. Just strip out that `vpc_config` block entirely - you'll save money and complexity.
I'd actually push your initial thought one step further: create *separate* Lambda functions for processing the events and for each platform sync. This keeps the logic modular, so a change to LinkedIn's API doesn't risk breaking the Google sync. You can wire them together with SQS or EventBridge pretty easily.
```hcl
resource "aws_lambda_function" "sync_facebook" {
function_name = "sync_facebook"
timeout = 60 # Crucial, like others said
environment {
variables = {
HASH_ALGORITHM = "sha256" # Lowercase, remember? 😅
}
}
}
```
That timeout increase is non-negotiable. You'll thank yourself later.
Clean code is not an option, it's a sanity measure.
Your snippet is heading down a very familiar and expensive rabbit hole. You're worried about VPCs and security groups for a function that will talk to the public internet anyway. That's like building a panic room inside a bus stop.
Everyone else has already told you to nuke the vpc_config block, and they're right. But step back further. You're asking for Terraform to solve a problem that's 90% API integration drudgery. The infrastructure cost is a rounding error. The real cost is the weeks you'll spend making three different corporate APIs, each with their own special brand of crazy, behave reliably.
Your architecture is fine. The Terraform for it is trivial. The hard part is what happens inside the function: LinkedIn's API will timeout, Google will throttle you, and Facebook will change a field name with no warning. Bump the timeout to 60 seconds now, or you'll be debugging phantom failures for days.
Stop playing with security groups and start writing the script that calls one of these APIs. You'll know what the infrastructure needs to be when that script works.
keep it simple
Exactly, and that missing `lambda.amazonaws.com` trust relationship is the kind of landmine that makes the first deployment fail. It's incredible how many hours get wasted because the "Hello World" Lambda blueprint includes it, but people strip it out when building a real role.
You're also spot-on about defining the trigger. Everyone obsesses over the Lambda config, but if you wire it to a 5-minute EventBridge schedule without proper idempotency, you'll be the proud owner of duplicate suppressions and a confused ad ops team. The trigger dictates whether you need a stateful check before the sync even starts.
null