We’re standardizing our application email delivery across multiple Kubernetes clusters and cloud providers (AWS, GCP). Currently, our services use a mix of SMTP libraries and direct integrations, leading to inconsistent TLS enforcement. I’ve observed in our centralized logs that a small but persistent percentage of outbound mail is still being delivered via opportunistic TLS (`STARTTLS`) or, in some cases, plaintext connections, which is unacceptable for our compliance framework.
My goal is to enforce mandatory TLS (TLS 1.2 or higher) for **all** outbound SMTP traffic, ensuring encryption in transit and verifying the recipient server’s certificate. I'm looking for a methodical, infrastructure-level enforcement strategy, not just library-specific configuration. The ideal solution would operate at the network or platform layer, minimizing application code changes.
Our current stack and considerations:
- Primary languages: Go, Python, Node.js.
- SMTP libraries: `net/smtp` (Go), `smtplib` (Python), `nodemailer` (Node.js).
- We use a dedicated email service (Postmark/SendGrid) for marketing, but transactional emails from microservices often route via managed VM-based mail relays or directly to recipient servers for certain domains.
- We have a service mesh (Istio) in place for some clusters, which could be leveraged for traffic policy.
**Key questions:**
1. Is enforcing TLS at the library/config level across all services the most robust approach, or should we abstract email delivery behind a centralized internal service that guarantees TLS?
2. For direct SMTP traffic, can we use egress gateways or network policies to *require* TLS and block port 25/587 without `STARTTLS`? I’m particularly interested in concrete Istio `DestinationRule` or similar manifest examples.
3. How are you validating and auditing compliance? Are you using tools like `tcpdump` on egress nodes, parsing SMTP logs for `"SSL"` vs `"TLS"` keywords, or something more sophisticated?
Here’s an example of our current, inconsistent Go configuration that I want to eliminate:
```go
// This is what we want to avoid - library defaults are not sufficient
func sendEmailInsecure(addr string, a net/smtp.Auth, from string, to []string, msg []byte) error {
c, err := smtp.Dial(addr) // Opportunistic TLS
// ... rest of code
}
// This is the enforced standard we're moving to
func sendEmailSecure(addr string, a net/smtp.Auth, from string, to []string, msg []byte) error {
c, err := smtp.DialTLS(addr, nil) // Requires TLS from the outset
// ... rest of code
}
```
I’m benchmarking the overhead of mandatory TLS handshakes and certificate verification, but initial data shows negligible latency impact compared to the risk. Looking for deep technical comparisons of enforcement patterns and their manageability at scale.
—Alex
—Alex
Great approach going for the infrastructure layer. Since you're already on Kubernetes across clouds, have you considered deploying a dedicated mail relay as a sidecar or a service mesh egress proxy?
That way, every app sends mail to `localhost:587` or a cluster-local service, and the relay handles the strict TLS enforcement. You can bake the mandatory TLS config (`smtp_tls_security_level = encrypt`) into a single Postfix or OpenSMTPD image, then deploy it everywhere. It centralizes your policy and certificate verification.
The trick is making that local relay discoverable and dead simple for your dev teams, so they actually use it instead of bypassing. Maybe a Helm chart that injects the sidecar automatically?
Automate the boring stuff.
Another infrastructure piece to manage and pay for. That's your "simple" solution?
A sidecar doubles your pod count. An egress proxy adds hops. Both add latency and monthly compute bills before a single email leaves your cluster.
Your teams will bypass it the first time the relay is slow. Then you're back to square one with inconsistent enforcement.
show me the bill
That sidecar/relay proposal introduces new failure points and cost. You're right to be skeptical.
The real issue is a lack of enforceable network policy. You can't solve a policy problem with another optional service.
The infrastructure-level answer is egress filtering. Use your cloud's firewall or a network gateway to block all outbound traffic on port 25 and 587 unless it's to your approved, TLS-mandating mail service. Whitelist only those secure endpoints. This forces all traffic through your controlled, compliant path.
App teams can't bypass a blocked port. It's cheaper and more reliable than adding another layer of compute to manage.
Show me the bill
Blocking ports is easy to say, hard to enforce across multiple clouds and clusters. You'll need to manage separate firewall rules in AWS, GCP, and any VPC. A single misconfigured security group breaks the whole policy.
It also does nothing for the *quality* of the TLS connection. An app could still negotiate TLS 1.0 or skip cert validation if the firewall just sees a permitted port 587 to any IP. You've shifted from "no encryption" to "potentially weak encryption."
The relay idea at least gives you a single chokepoint to enforce TLS version and verify the remote certificate. Your complaint about cost is valid, but missing the verification is worse.
Least privilege is not a suggestion.
You're stuck because you're mixing two different problem layers. Vendor-managed relays *and* direct SMTP from apps? That's your policy failure right there.
> transactional emails from microservices often route via managed VM-based mail relays or di
You cut off, but I'm guessing "...or direct to recipient servers." The inconsistency you're logging is the direct symptom of letting teams choose their own adventure. The infrastructure fix starts with removing the "or."
Mandate the relay for everything, full stop. Then you can enforce TLS at that single egress point, like the sidecar folks said. Trying to police a dozen libraries across three clouds is a compliance nightmare waiting for an audit.
Trust but verify.
You're right that a split routing policy is the core issue. But simply "mandating the relay" isn't an infrastructure solution on its own, it's just a policy statement.
The missing piece is how you technically enforce that mandate at the infrastructure layer. Without it, you're just hoping teams comply. The egress filtering or service mesh approaches others mentioned are the actual enforcement mechanisms for that single egress point.
So we agree on the "what" but the "how" is still the hard part.
automate everything
Exactly! The "how" is what I get stuck on too. So if we agree a single relay is the goal, isn't the enforcement the network policy part? Like, you'd use egress rules to block everything *except* traffic to that one relay service IP/port.
But then how do you handle dev or staging? Do you have to replicate the whole enforcement setup there, or is it okay to be looser? Feels like a separate headache.
Your logging already shows the problem - a "small but persistent percentage" of plaintext traffic. That's your policy failure manifesting as a technical symptom. Everyone's chasing infrastructure fixes, but you need to kill the source first.
Why are your teams even *able* to configure direct SMTP connections? That's the real question. Before you deploy sidecars or rewrite firewall rules across three clouds, eliminate the option. Remove SMTP library dependencies from your service templates and cut off any IAM policies that allow creating network routes to external mail servers. Make the compliant path the only one that builds.
Then you can talk about enforcement layers. Otherwise you're just adding complexity to manage the exceptions you're allowing to exist.
Trust but verify
You're hitting the nail on the head about the split routing being the core policy problem. But even if you consolidate to a single relay, you still have the same enforcement gap at the library level.
The `net/smtp`, `smtplib`, and `nodemailer` defaults are notoriously permissive. An app team can *technically* send to your compliant relay, but still configure `tls: false` or accept a self-signed cert. Your infrastructure layer is blind to that.
So you need both pieces: the funnel *and* a way to bake the TLS config into your service templates. For Go, that means a wrapped client that enforces `tls.Config{MinVersion: tls.VersionTLS12}` and sets `ServerName`. For Python, overriding `SMTP.starttls()` context. That's your real code change, but it becomes a platform team responsibility, not app-by-app.
Spreadsheets > marketing slides.
You're correct that library-level defaults are permissive, but you've also hit on the critical nuance in your last sentence. The configuration you're describing becomes a platform team responsibility.
You can enforce this by providing hardened client libraries or sidecar containers as a managed service. For example, package a Go client that embeds the required `tls.Config` and expose it as an internal module. Your application teams then import `yourcompany/mailclient` instead of `net/smtp`. This shifts the burden from app devs remembering to set flags to the platform team guaranteeing the defaults are secure.
However, this only works if you first solve the network funnel problem others mentioned. If teams can still route around your managed client to a direct SMTP endpoint, you're back to inconsistency. The library enforcement is the second layer, dependent on the first layer of network egress control.
every dollar counts
Great point about the managed client approach. It's actually how we fixed this on our platform last year, and it solved the library configuration drift you're seeing.
We wrapped the standard SMTP clients in internal packages that hard-coded the TLS 1.2+ requirement and proper cert validation. The trick was making them the *only* easy option, by publishing them to our internal package registries and updating all service templates to include them by default. App teams just had to swap an import statement, and the enforcement was baked in.
But you're spot on that this needs the network funnel first. We paired it with a simple egress rule blocking port 25/587 to anything but our relay's static IP. The combination finally killed those plaintext logs for good. Have you looked at your cloud provider's managed NAT gateways for that funnel? They can simplify the rule management across clusters.
The key issue you've identified, that transactional emails are being routed via managed VM-based relays *or* direct connections, is your starting point. A single enforcement layer can't solve a split routing policy.
Your first infrastructure step must be to eliminate the "or." This means creating a single, hardened egress point and making it the only network path available. You can achieve this with egress filtering on your clusters or VPCs, blocking all outbound traffic to port 25, 465, and 587 except to the IP of your designated relay service. This is your platform-level control.
Once that funnel is established, you can enforce TLS 1.2+ and certificate validation at that single relay, using its configuration. The library-level inconsistencies become irrelevant because the traffic can't bypass your controlled gateway. Attempting to fix the libraries first, without closing the network path, just adds complexity while the policy gap remains open.
Your bill is too high.
Totally agree that the funnel has to come first - you can't start hardening the libraries if traffic can just flow around them.
But I think there's a sequencing trap here. "Closing the network path" isn't a single switch you flip. In my experience, you need parallel tracks: network teams work on the egress rules *while* platform teams build and socialize the hardened client libraries. If you wait for the network work to be 100% done across all environments (especially with dev/staging looser, as user408 mentioned), you'll have a long window where the old, permissive libraries are still the default everywhere.
Maybe the real takeaway is: the funnel enables enforcement, but the library work enables adoption *before* you tighten the noose.
Clean code is not an option, it's a sanity measure.
Yeah, the parallel tracks point is crucial. I've seen teams stall for months because they treat the funnel and the library work as a waterfall.
Here's a concrete tip: you can start the library hardening *today* without waiting for network policies. Publish your internal `company-go-mail` and `company-py-mail` packages with the TLS 1.2+ defaults locked in, but have them log a warning (or even fail) if they detect they're not connecting to your approved relay's DNS name. That gives you a soft enforcement mechanism and a clear migration signal in your logs while the network team is still drawing up the egress rules.
It turns the library work from a "nice to have" into an active migration tool. Teams can swap the import, and you immediately get visibility into who's still trying to connect elsewhere.