Just spent the last six weeks navigating the "seamless integration" between Cato SASE and our existing on-prem SD-WAN (a big-name vendor, let's call it Vendor X). The sales pitch promised a harmonious blend. The reality, as usual, involved a surprising amount of duct tape and some creative interpretations of "failover."
We now have a setup where our primary tunnels terminate to Cato PoPs, with Vendor X's tunnels as a backup path, triggered by BGP metrics and some event-driven scripting Cato insists on calling "orchestration." The fun part was getting the two systems to agree on a health check. One expects ICMP, the other wants a TCP handshake to a specific service port, and neither would accept the other's definition of "unhealthy."
So, it's working. Traffic fails over, the CFO is happy because we're "multi-cloud," and I'm left wondering if the management overhead is worth the theoretical resilience. Ask me anything about the gotchas, the config quirks, or the inevitable billing surprises that come with bolting one managed service onto another.
/c
Beware of free tiers
Oh, the health check negotiation. That's the real "integration" step they don't put in the datasheet. We had a similar dance between two different platforms last year, and ended up having to deploy a tiny, purpose-built container just to answer both probe types. It works, but it felt like building a translator for two systems that claim to speak the same language.
Was the BGP metric tuning the main lever for failover timing, or did you have to lean heavily on that event scripting to get an acceptable convergence time? I've found the scripting can sometimes add its own latency spikes.
Ship fast, measure faster.
That translator container idea is actually pretty clever, I might have to borrow that. 😄
We mostly relied on BGP tuning to keep it simple. The event scripts felt like another potential point of failure, so we kept them to a bare minimum, just for some specific service alerts. The latency was okay for us, but you're right to be wary.
Did you notice any extra overhead from running that container during normal ops, or was it negligible?