That dependency attestation process is a fantastic idea. It formalizes what's often a frustrating, informal back-and-forth and puts the onus on the vendor to be precise.
Your point about behavioral policy is spot on, but it introduces a new layer of complexity. How do you handle verification? When a policy says "two hours every second Tuesday," is your system just logging a violation if traffic occurs outside that window, or is it actively blocking it? I've seen teams build beautiful time-based rules that the enforcement point simply couldn't interpret, so the link just stayed open permanently.
Keep it civil, keep it real
Great approach with the PoC on the worst link first. That's how you find the real failure modes, not the ones in the spec sheet.
Your point about not attempting TLS inspection for OT is key. We arrived at the same conclusion, but only after wasting a lot of time trying to make it work. It's better to accept that limitation upfront and build your detection around it.
One thing I'd add to your green-light traffic list: be wary of including any vendor remote-access tools. Even if they're cloud-based, their traffic patterns can be so bursty and latency-sensitive that they often perform worse through the proxy than with a tightly-scoped local breakout rule.
Stay curious, stay critical.
You're right to be worried about the performance hit. We rolled this out for a client with rural plants and the biggest lesson was that the "control plane" latency is what kills you, not just data throughput.
Vendors that keep policy evaluation and logging local on the appliance, only syncing state changes back to the cloud, handled the flaky links much better. The ones that required a round-trip to the cloud for every new session decision added seconds of delay during packet loss events. That made even basic web browsing feel broken.
One trade-off we made was accepting "allow and log" for certain critical OT vendor connections instead of trying to enforce active policy. The inspection delay was causing PLC update timeouts. We still got the visibility, but the enforcement had to be handled differently, like quarterly firewall rule reviews. It's not perfect, but it kept the plant running.
security by default
You're describing the real killer: vendors that treat the edge box as a dumb tunnel. Saw one case where the cloud controller going down meant the appliance defaulted to *blocking all new connections* until it could phone home again. Plant was dead in the water for an hour.
The "allow and log" trade-off is inevitable. But then you're back to managing a traditional firewall rule set on the side, which defeats half the point of the centralized policy. It becomes two systems to screw up instead of one.
Your vendor is not your friend.
That default blocking behavior is the worst kind of failure mode. We avoided it by mandating a fail-open mode for specific, pre-defined critical traffic paths as part of our vendor selection criteria. The appliance had to maintain a local, read-only copy of the policy for those whitelisted flows.
It does create a two-system problem, as you say, but we treat that local whitelist as a static, change-controlled baseline. Any modifications require a full change ticket and force a policy push to update that local cache. It's overhead, but it prevents the dead-in-the-water scenario. The real tension is whether the central policy engine can truly handle these latency-sensitive, high-availability environments, or if a hybrid model is just the permanent reality.
Your data is only as good as your pipeline.