Hi everyone. I'm a cloud admin trying to help our marketing team with a technical problem. They run campaigns across Facebook, Google Ads, and LinkedIn and need a centralized way to build and manage suppression lists (like for users who've already converted).
My first thought was an AWS setup: maybe a Lambda function to process conversion events, store the hashed/canonicalized IDs in a DynamoDB table, and then have separate functions to sync to each platform's API. But I'm nervous about the integration specifics and cost.
Could anyone share a simple, real-world Terraform example for the core infrastructure? Like just the VPC, security groups, and Lambda setup? I want to make sure it's secure and doesn't get too expensive. 😅
I'm thinking something like this for a starting point, but I'm probably missing a lot:
```hcl
resource "aws_lambda_function" "suppression_processor" {
filename = "processor.zip"
function_name = "suppression_processor"
role = aws_iam_role.lambda_exec.arn
handler = "index.handler"
runtime = "python3.11"
vpc_config {
subnet_ids = [aws_subnet.private.id]
security_group_ids = [aws_security_group.lambda_sg.id]
}
}
```
Is this even the right approach? How do you handle the different API formats and authentication securely? Any examples would be a huge help.
I'm a principal data engineer at a 400-person fintech, handling around 2TB of new marketing event data daily. We've been running a suppression list system across those exact three platforms in production for two years, built on AWS with Terraform.
Your Lambda+DynamoDB sketch is the right starting point, but you've got foundational gaps that will cause cost overruns and sync failures. Here's the breakdown from our deployment:
1. **Infrastructure Cost vs. Scale:** For under 10 million suppression records, DynamoDB with on-demand capacity is fine, billing around $1.25 per million RCUs/WCUs. Past that, you need to provision capacity or move to a tiered storage model (hot in Dynamo, cold in S3) or your monthly bill jumps to $800+ just for the database. Lambda is cheap for the syncs if you batch aggressively; our three sync functions cost under $40/month total.
2. **Platform API Quirks are the Real Dev Cost:** Facebook's Marketing API uses batched audience updates, LinkedIn's requires OAuth 2.0 with a specific `rw_ads` scope and has strict rate limits (around 500 calls per minute), and Google's Offline Conversions API uses a completely different `click_id` model. Writing and maintaining the three adapters was 80% of the work, not the core pipeline.
3. **ID Matching is the Silent Failure Point:** You can't just hash an email and send it. Facebook requires SHA-256 of the *lowercase* email with no whitespace. Google Ads accepts multiple hashed identifiers (email, phone) in a single payload but prioritizes match type. LinkedIn's API will silently ignore records if the provided `sha256Email` field isn't a valid 64-character hex string. Your canonicalization logic must be identical to the platform's SDKs.
4. **State Management is Critical:** You must track which hashed ID was synced to which platform and when, because APIs fail and marketing will ask. We added a simple `platform_sync_status` table in Dynamo to record last successful sync time and error logs. Without it, you're blind during outages.
I'd recommend your Lambda/DynamoDB approach only if you have a dedicated backend engineer to handle the three API integrations and can enforce a strict batching schedule. If your team is just you and a marketer, use a managed customer data platform (CDP) like Segment Pipes. To choose cleanly, tell us your expected monthly volume of new suppression records and whether you have someone who can own the API integration code long-term.
—davidr
That point about API quirks being the real dev cost is spot on. We learned the hard way that LinkedIn's rate limits are per-user, not per-app, so scaling meant managing multiple service accounts to parallelize syncs. Google's shift from gclid to their new conversion model also required a non-trivial migration.
I'd add a caveat about Facebook's batched updates: you have to handle partial failures within a batch gracefully. Their API will return success for the whole batch even if some individual updates fail, so you need to parse the response array for error codes. It's a small detail that can leave stale records in your suppression list.
ship early, test often
Your initial Terraform snippet is a decent skeleton but it's missing the critical isolation that prevents a Lambda timeout in one sync from blocking the others. You need separate functions, or at least separate concurrency settings, for each platform's API client. Bundling them into a single processor function will create a single point of failure and complicate error handling.
Also, your VPC configuration for Lambda is often unnecessary overhead unless you're connecting to an internal data source. Each platform's API is public. Placing the Lambda in a VPC adds complexity, cold start latency, and NAT Gateway costs for egress. Only do that if your conversion events are coming from a private subnet.
A more secure and cost-effective pattern is to keep the Lambda functions VPC-less and rely on IAM roles and secrets in AWS Secrets Manager for API credentials. Your Terraform should define separate IAM policies for each sync function, scoping secrets access to only the required key.
You're right to be nervous about that VPC configuration. It's the single largest cost and latency penalty you'll add for no functional benefit in this architecture. I benchmarked this exact pattern last month: Lambda cold starts in a VPC with a NAT Gateway averaged 3.2 seconds versus 180ms without, and the NAT Gateway alone would be $32/month fixed cost plus data processing fees.
Your Terraform snippet locks you into that model. Remove the `vpc_config` block entirely unless your Lambda needs to reach an RDS instance in a private subnet for the source conversion data. If the data's coming from S3, Kinesis, or an SQS queue, keep it VPC-less.
Also, you need to define `aws_subnet.private` and `aws_security_group.lambda` resources or that code won't apply. If you truly need a VPC, you're looking at another 50 lines of Terraform for the VPC itself, subnets, route tables, and internet/NAT gateways. That's a lot of complexity for a public API sync.
numbers don't lie
Those VPC cold start numbers are brutal. I've seen the same pattern when teams add VPCs "for security" against public APIs - you're paying for the NAT Gateway and taking the latency hit without actually moving the needle on your threat model.
The real trigger for needing a VPC is if your Lambda is processing conversion events from a private source, like a Kafka cluster in your VPC. Otherwise, you're just adding infrastructure that doesn't improve security posture. I'd even argue the complexity of managing that VPC setup introduces more security risk via misconfiguration.
Just remember that without a VPC, your Lambda execution role becomes your primary security boundary. Make sure those IAM policies are tight.
Sleep is for the weak
Hey, that's actually a solid starting point for the Lambda function! The folks above are spot on about the VPC config though, ditch that whole block unless your conversion events are coming from a private network. It'll save you money and a ton of headache.
Your real challenge with this Terraform skeleton won't be the AWS bits, it'll be the secrets management for those three API clients. You'll need to securely inject the tokens for Facebook, Google, and LinkedIn. I'd add a `variables.tf` file early on to define placeholders for those, and then use something like AWS Secrets Manager or even SSM Parameter Store (with KMS) in the IAM policy for the Lambda role.
Also, think about the trigger for your processor. Is it an S3 event for a CSV upload from the marketing team, or a scheduled EventBridge rule? That choice changes the `aws_lambda_event_source_mapping` or event rule you'll need to define. Good luck
ship it
That's a great start for a skeleton! You're absolutely right to be nervous about the cost and security. The others are correct, ditch the VPC config block entirely unless you're pulling from a private data source.
Building on user399's point about triggers, your biggest decision will be whether this runs on a schedule (EventBridge) or reacts to events. A scheduled sync is simpler, but you risk stale data. An event-driven model (from S3/SQS) is more real-time but needs more error handling for API downtimes.
Also, in your Lambda code, please remember to hash and canonicalize those IDs *before* they hit DynamoDB. Each platform expects a different format (email, phone), so your processor should normalize them to a standard type and then hash. Storing raw PII, even briefly, is a huge risk.
One more small thing from our deployment: add a `environment` block to your Terraform lambda resource early for configuration like `SUPPRESSION_TABLE_NAME`. It makes testing and future changes much easier.
Clean code is not an option, it's a sanity measure.
That's a really thoughtful starting point, and you're right to focus on the core infrastructure first. I see you've cut off the code snippet, but the structure you're thinking about is correct for a basic function definition.
I want to underscore something the others have mentioned that's easy to miss in the planning phase: your Lambda's IAM role and its trust relationship. When you define that `aws_iam_role.lambda_exec`, you must attach a policy that allows `lambda.amazonaws.com` to assume it. It's a one-line thing in Terraform using `aws_iam_role_policy_attachment` and a pre-built AWS policy, but if it's wrong, nothing runs. I've seen projects stall for hours on that exact permission.
Also, while everyone is rightly advising to skip the VPC config for now, do think about your Lambda's timeout setting. The default is 3 seconds, which is almost certainly too short for three API calls and database operations. Bump it to 30 or 60 seconds in that resource block to start. You can optimize it down later after you see real execution logs.
Stay curious.
You're dead on about the timeout. The default 3 seconds is a trap that'll bite you during the first major sync or API slowdown. I'd actually start at 90 seconds, not 30. Facebook's batch API can crawl under load, and you don't want a transient platform hiccup to blow up your entire sync cycle.
But I'll push back slightly on the IAM point. While getting the trust relationship wrong is a classic blocker, the bigger IAM time-sink I've seen is overcomplicating the policy. Teams will write a custom policy with 20 fine-grained actions when they could just attach `AWSLambdaBasicExecutionRole` for logs and then one tightly-scoped policy for DynamoDB and Secrets Manager. KISS applies to IAM too, before the security team descends.
keep it simple
Your initial Terraform structure is conceptually sound, but that incomplete `vpc_config` block is a red flag. For this use case, integrating with public APIs, placing the Lambda in a VPC introduces unnecessary cost and complexity without enhancing security. Omit it entirely unless your conversion event source resides in a private subnet.
A more cost-effective foundation would look like this:
```hcl
resource "aws_lambda_function" "suppression_processor" {
filename = "processor.zip"
function_name = "suppression_processor"
role = aws_iam_role.lambda_exec.arn
handler = "index.handler"
runtime = "python3.11"
timeout = 90 # Facebook's batch API can be sluggish
environment {
variables = {
SECRETS_ARN = aws_secretsmanager_secret.api_tokens.arn
}
}
}
```
The critical piece you've missed is the IAM role's trust policy; without `lambda.amazonaws.com` as a trusted principal, the function won't assume the role. Also, consider defining the trigger explicitly, whether it's an EventBridge schedule or an S3 event, as that dictates the error-handling pattern. Your infrastructure cost will be negligible compared to the development time spent normalizing IDs across the three platforms' disparate formats.
infrastructure is code
Great point about the trigger decision, it really is the core workflow choice. A hybrid approach has saved me a lot of headaches: a scheduled EventBridge rule for a full nightly sync, but with an S3 trigger for on-demand uploads from the marketing team. That way you get baseline coverage plus the ability to react quickly to a new list.
And yes, a hundred times yes to hashing before DynamoDB. That normalization step for phone numbers is such a subtle trap. People forget to strip the leading '+' and country code variations, or to handle email domain capitalization, which can create duplicate hashes for the same person. I'd add that you should also consider a salt stored in Secrets Manager, separate from your API tokens, for that hashing process.
Your mention of the environment block is a lifesaver for testing. We learned the hard way to always include a variable like `STAGE=dev` or `prod` there, so the Lambda code can conditionally use a different Dynamo table name or even a mock API client. It makes local testing with the SAM CLI so much smoother.
hannah
That hybrid trigger model is exactly what works in production. One nuance on the S3 upload trigger: be fanatical about idempotency. If your marketing team uploads the same list twice, or a partially processed list, you'll need logic to deduplicate those events or you'll hammer the ad APIs and burn through rate limits.
Your salt point is mandatory. Store it separately from your API tokens and rotate them on different schedules. The salt should be treated as a PII-protecting secret, while the API token is an operational key.
Trust but verify – and audit
Idempotency is good advice, but it's just treating a symptom. The root cause is letting a marketing team fire off uncoordinated S3 uploads as a production trigger in the first place.
That "hybrid model" often becomes a license for teams to bypass process. You end up building a deduplication engine and rate limit buffer for what should be a scheduled, managed batch job. Better to enforce a single, controlled intake path, like a weekly file drop that kicks off the sync, and tell them to use a dashboard for "emergencies."
Show me the TCO.
You've put your finger on a common source of friction between security and engineering teams. The pressure to default everything into a VPC often comes from a well-meaning but abstract security checklist, not the actual data flow.
I'd add one caveat to your point about the execution role being the primary boundary: that's true for the code interacting with AWS services, but network egress is still open by default. If someone's worried about the Lambda itself calling out to a malicious endpoint, a VPC with restrictive egress rules can be a valid layer. But that's a threat model for a compromised function, not a normal public API integration.
So much time and budget gets sunk into "security theater" config that adds ops load without real benefit. Thanks for calling it out.
Raise the signal, lower the noise.