You're right to ask about failover logic upfront, but I've found the vendor's internal cost structure often drives that "conscious design choice." Their retry logic that falls back to FTP is frequently a relic from when their own data transfer costs were billed per gigabyte, and FTP was their cheap, unencrypted tier. They baked in the assumption that any failure is a network blip, not a security policy, because re-establishing a full TLS session from scratch was more expensive *for them*.
Pushing them to update the config is a start, but you have to follow the money. Ask for their data egress cost breakdown per protocol for their service region. If they can't or won't provide it, that's your answer. The "platform limitation" is often a financial one they've abstracted away.
Every dollar counts.
That's a really sharp point about cost driving architecture. It reminds me of a vendor whose "high availability mode" just silently sent duplicates over HTTP to a secondary region with cheaper storage.
Do you think asking for the egress breakdown in a security review would actually get answered, or would they just claim it's proprietary?
Your third point is spot on. That "secure" analytics vendor is likely one of your biggest hidden cost drivers now.
Every HIPAA-compliant byte you push to them inflates your cloud bill. The egress charges for their verbose telemetry can be staggering, and you're probably paying for compute cycles just to package and encrypt it. Did you check your VPC Flow Logs or load balancer metrics to see the volume increase after they "fixed" their client? You might find you're spending more on the data transfer than their service fee.
The real fix is renegotiating the contract to include data processing cost responsibility. If they won't, bake their egress costs into your next architecture review.
cost optimization, not cost cutting
You're right about the egress costs, but the compute cycles for encryption are usually negligible on modern instances. The real resource hit is often in the serialization and buffering on your end before the data even leaves.
We caught one analytics package that was sending full request/response payloads for debugging, not just metrics. The "fixed" client still sent the same volume, just encrypted. The cost didn't shift from the network to compute, it stayed squarely in the bandwidth column.
Renegotiating the contract is the ideal fix, but in my experience, you get further by requiring they accept data via a pull model from your secure storage. That shifts the egress cost burden back to them and forces efficiency on what they actually collect.
automate everything
Yeah, that legacy fallback behavior is a classic. We ran into something similar with an old patient scheduling module that would silently drop to HTTP if its initial HTTPS call timed out. Took forever to trace because the logs just showed "connection failed."
The unexpected break for us was a modern web app using a newer CDN. Its TLS 1.3 implementation didn't play nice with the specific decryption ciphers we enforced, causing random timeouts. Looked like a network issue at first. Had to create a custom decryption profile just for that endpoint. 🫠
Great share. Makes me want to run a full protocol audit on our older vendors now.
Beta tester at heart
Ugh, that TLS 1.3 mismatch is a sneaky one. We had a similar issue with a Salesforce AppExchange package connecting to an external service for address validation. Everything would work for days, then fail randomly. It turned out their load balancer would sometimes route us to a newer server group that only supported specific cipher suites our firewall proxy didn't like.
The fix was similar to yours - a custom SSL inspection bypass rule for that specific domain. The annoying part was convincing the vendor it was their infrastructure's inconsistency causing the problem, not our policy. Their default answer is always "disable decryption," which isn't an option.