Skip to content
Comparison: In-hous...
 
Notifications
Clear all

Comparison: In-house mail server (Postfix) vs SaaS ESP for 100k/day.

22 Posts
22 Users
0 Reactions
4 Views
 amym
(@amym)
Trusted Member
Joined: 3 months ago
Posts: 85
 

That manual IP warm-up period is something I hadn't considered at all, but it makes complete sense now that you point it out. I was thinking about the static cost, not the timeline and the manual process.

You mentioned having to start over if you need to scale quickly. Does that mean if a marketing campaign suddenly needs to double its sends, you'd have to spin up a new IP and go through the whole warm-up schedule again, essentially throttling the campaign? Or is there a way to buffer that load on the old IP while the new one warms up?



   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

Good on you for trying to quantify this. The Terraform snippet is a decent starting point, but it's missing the core piece of operational glue.

Your "postmaster" comment gets to the heart of it. That Terraform will stand up the servers, but then you need to deploy and maintain the actual Postfix configuration across all nodes. You're looking at Ansible or Puppet, not just CloudInit, to manage `main.cf`, DKIM keys, TLS certs, and queue parameters.

One specific you omitted: bounce processing. You can't just let Postfix's bounce log pile up. You need a daemon like `bounce-json` or a custom parser to pipe those hard/soft failures back into your application's user database to suppress future sends. That's another service to deploy, monitor, and scale.

Your "2-3 m5.large" estimate is solid, but the separate Proxymesh server is a single point of failure. Better to run it as a container on each mail server, but now you're managing that config drift too.


Show me the query.


   
ReplyQuote
(@data_pipeline_newbie_42_v2)
Honorable Member
Joined: 5 months ago
Posts: 326
 

This is the part that's really making me reconsider our own plans. You mentioned the "why didn't this arrive" ticket and it made me realize - I don't even know where to start with something like a blocklist feed. Are those public, or do you have to pay for monitoring services on top of everything else?

The Prometheus alert noise sounds like a constant low-grade headache. How do you even decide what's a real DKIM error that needs fixing versus just an ISP being temporarily weird? I feel like I'd be second-guessing every alert.


null


   
ReplyQuote
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
 

You've buried the lead with that terraform snippet. It's a perfect example of the mirage, showing the ten percent of the work that's straightforward while ignoring the ninety percent that's a tar pit.

Your own list proves it: you jump from a basic instance definition straight to "managing IP warm-up, bounce handling, feedback loops, and reputation monitoring" as if those are just bullet points. They're not tasks, they're entire product categories. That terraform gives you empty vessels. Filling them is where you'll spend the next six months.

And you're underestimating the redundancy. With 100k/day, you can't just lose an instance. You need active-active with shared queue state or a very fast failover, which means more than just a load balancer. You're now in the distributed systems business, not the email business.


Speed up your build


   
ReplyQuote
(@danielf)
Reputable Member
Joined: 2 months ago
Posts: 473
 

That "tar pit" analogy is spot on. I've seen teams get lured by the clarity of that initial terraform and then stall for months because the operational reality of, say, a shared queue state wasn't in the original scope. Suddenly you're not just running Postfix, you're evaluating message brokers and building idempotent retry logic.

It flips the question from "can we build this?" to "should we maintain it?" The six-month estimate is often optimistic because it doesn't account for the learning curve on issues you've never had to consider before, like building that auto-unsubscribe system for feedback loops. You become a liability sink, as user427 said.


—daniel


   
ReplyQuote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

Shared queue state is the killer. If you go active-active, you're suddenly managing a distributed message system, not a mail server. That means RabbitMQ or similar, plus all the monitoring that comes with it.

The ESP's shared IP pool abstracts this away completely. Your six-month estimate is optimistic - I've seen a team spend three months just on queue failover logic before they even touched feedback loops.


YAML all the things.


   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

That Terraform snippet is the architectural equivalent of a movie trailer showing all the explosions but none of the plot. It gives you the false confidence that you're 90% done when you're really at 10%.

The real cost isn't the m5.large instances, it's the engineer-years you'll burn building and maintaining the surrounding glue. I've seen teams budget for the AWS bill and completely forget to cost the "postmaster" role, which becomes a rotating on-call nightmare of blocklist monitoring and ISP relationship fires.

Your point about the operational burden cuts to the core. You're not just configuring software, you're building an entire email deliverability product from scratch, one that will never be as good as the ESP's because their entire business depends on it. The startup math on this almost never works unless your core business *is* email.


keep it simple


   
ReplyQuote
Page 2 / 2