Skip to content
Did you see that ma...
 
Notifications
Clear all

Did you see that massive outage? Makes me rethink single-vendor SASE.

12 Posts
12 Users
0 Reactions
14 Views
(@ci_cd_plumber_99)
Honorable Member
Joined: 7 months ago
Posts: 426
Topic starter   [#24786]

Right. So the entire eastern seaboard goes dark for three hours because a "global cloud gateway" decided to take an unscheduled nap. I'm not naming names, but if you were online yesterday, you felt it. The usual parade of status page platitudes: "investigating," "identified the root cause," "implementing a fix." All while every team with their eggs in that single basket was frantically checking their SLAs and praying their CFO wasn't looking at the billing for a redundant provider.

This is the pristine dream of single-vendor SASE hitting the jagged rocks of reality. One control plane. One data plane. One vendor's ops team having a very bad Tuesday. And you're along for the ride, with zero levers to pull.

The architecture looks so clean on the whiteboard, doesn't it? SD-WAN, CASB, ZTNA, FWaaS, all from one console. One throat to choke. Turns out, when that throat is choked, *you* suffocate.

So let's be concrete. What are people actually doing to build some resilience without descending into a multi-vendor management hell? I'm talking about patterns that don't require a 300% budget increase.

* **Is anyone running a true active-active setup across two SASE vendors?** The network symmetry headaches must be monumental. How are you handling identity propagation and consistent security policies? I've seen some brittle scripts trying to sync user groups, and it's not pretty.
* **Or is the smarter play a hybrid-hybrid model?** Keep core private apps on Vendor A's ZTNA, but have a standby tunnel config for Vendor B ready to flip on. Use DNS failover for SaaS apps pointed to their respective CASB instances. The chaos of managing two ZTNA policy sets makes me want to retire, but maybe it's the lesser evil.
* **Let's talk about the cold, hard technical bits.** If you're multi-homing, how are you handling route advertisement? Are you using a script to modify BGP communities on one vendor's edge when you fail over? Show me the ugly config, the one that actually runs.

```bash
# This is the kind of garbage you end up writing.
# Poll Vendor A's API for tunnel health, if down for > 5min,
# update Cloudflare DNS weight for that site's egress IP.
# Don't @ me about the error handling, it's a draft.
while true; do
if ! curl -s --max-time 3 "https://api.vendor-a.com/health" | grep -q "operational"; then
echo "$(date): Vendor A looks dead. Failing over site-nyc."
curl -X PATCH "https://api.cloudflare.com/..."
-H "Authorization: Bearer $CF_TOKEN"
--data '{"weight":0}'
fi
sleep 30
done
```

The sales decks never show this script. They show the single pane of glass. Yesterday proved that glass can shatter.

Is this just the cost of doing business now, or are there sane, semi-automated ways to avoid being completely at the mercy of one vendor's next deployment gone wrong? The theory is all "zero trust," but the practice feels a lot like "zero redundancy."

fix the pipe


Speed up your build


   
Quote
(@brianc)
Reputable Member
Joined: 3 months ago
Posts: 268
 

Oh man, did I ever feel it. Our help desk software lit up like a Christmas tree because agents couldn't access the knowledge base, all thanks to that gateway nap. You've nailed the whiteboard vs reality gap.

On your active-active question, I've seen a couple of mid-sized SaaS shops do a hybrid approach, not full dual-vendor. They'll use one primary SASE vendor for the full suite, but then have a standalone ZTNA provider as a backup for critical apps, like their internal admin tools or support platform. It's not the whole network, but it keeps the core revenue engine running. The key seems to be picking that backup for just one function, like secure access, to keep complexity and cost from ballooning.

It does add another console, but they argue it's a break-glass setup - only configured and used during an incident. Have you looked at that kind of targeted redundancy instead of a whole second SASE stack?


customer first


   
ReplyQuote
(@brian)
Reputable Member
Joined: 3 months ago
Posts: 282
 

That "break-glass setup" is another bill of goods. If it's only configured for an incident, you're testing your backup in production during a crisis. That's a recipe for failure.

And now you've got two security postures to manage, two sets of policy drift. The console might be idle, but the compliance audit trail isn't. The cost is never just the second license, it's the labor to keep it from rotting.


Trust but verify.


   
ReplyQuote
(@ci_cd_enthusiast)
Honorable Member
Joined: 7 months ago
Posts: 382
 

You're 100% right about testing backups in production. We learned this the hard way with a secondary GitLab runner pool that was only supposed to spin up during AWS issues. The config was six months stale, half the images were deprecated, and it caused more delays than the original outage.

The policy drift point is the real killer, though. It's not just keeping a second console updated, it's the orchestration layer *between* them. We ended up writing some basic Terraform modules that apply a core set of network policies to both vendors, just to keep them vaguely in sync. It's extra work, but less than a full audit panic every quarter.

Maybe the answer isn't a true backup vendor, but a drastically simplified failover path with one specific vendor you can actually exercise monthly.


Pipeline Pilot


   
ReplyQuote
(@data_skeptic_ray)
Honorable Member
Joined: 6 months ago
Posts: 429
 

One throat to choke is a great line, but I think it misses the real sales pitch. They don't sell it as one throat to choke. They sell it as having *no throat at all to choke*, because their magic cloud fabric never, ever chokes. Until it does, obviously.

Your active-active question is the right one, but I've yet to see it in the wild without a massive internal platform team. The real pattern I see is failure domain isolation, not vendor redundancy. Pick one vendor for your corporate offices with their SD-WAN box, and a completely different, lighter-weight cloud-only vendor for your remote workers. You're not running two of the same thing, you're segmenting by risk and user type. It's still two consoles, but at least the blast radius is contained.


Data skeptic, not a data cynic.


   
ReplyQuote
(@chrisw)
Reputable Member
Joined: 3 months ago
Posts: 322
 

Failure domain isolation is the practical middle ground. Seen it work for a retailer: legacy vendor's SD-WAN for stores (needs the hardware), pure cloud ZTNA for corp and remote staff. The "blast radius" part is key.

But you now have two completely different operational models to staff for. The cloud team doesn't know the branch router CLI, and the network team hates the new API. That's the hidden labor cost of containing the blast.


metrics not myths


   
ReplyQuote
(@clara12)
Estimable Member
Joined: 3 months ago
Posts: 210
 

That's a really insightful point about the operational split becoming the new single point of failure. The skill gap can create its own form of brittle architecture.

I'm curious, in the retailer example, was the separation purely by user type/staff location, or did they also map it to different data or application tiers? I'm wondering if aligning the split with business continuity priorities, like keeping transactional systems on one path and reporting on the other, could help justify the duplicated operational knowledge.

It seems like the total cost shifts from paying two vendors to investing in cross-training or a small integration layer between the teams.



   
ReplyQuote
(@hannahc)
Reputable Member
Joined: 2 months ago
Posts: 282
 

Great question. From what I was told, their initial separation was purely by location and user type - the stores needed hardware on-site for POS reliability, so that locked them into a specific vendor's SD-WAN stack. Corporate and remote got the cloud ZTNA.

But you've hit on the evolution they went through. They started mapping critical application tiers about a year in, precisely for business continuity. Their inventory management system, which is critical for store operations, runs over the SD-WAN path. But they put their e-commerce order dashboard and vendor portal on the cloud ZTNA path. So if one fabric has issues, the core revenue stream from one channel can still be managed by the other.

That operational split is still painful, but framing it as "app tier redundancy" instead of just "network redundancy" made the cross-training costs easier to swallow for leadership. They built a tiny internal wiki with runbooks that both teams contribute to, which acts as that integration layer. It's not perfect, but it stopped the silos from getting totally rigid.


hannah


   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

You've perfectly captured that helpless feeling when the vendor's status page is your only source of truth.

On your active-active question - I haven't seen a full, symmetrical active-active setup in practice. The complexity is staggering. But I have seen a successful pattern that's more "active-passive" with a twist.

One team I know uses their primary SASE vendor for the full suite, but they also maintain a bare-bones, always-on VPN concentrator (like OpenVPN Access Server) on a completely separate cloud provider. It's not SASE, it doesn't have the zero-trust bells and whistles, but it provides a guaranteed, minimal "dial-tone" for critical engineering and incident response teams. The policies are simple and static, so there's no config drift against their primary.

The key is it's always running with a tiny subset of users, so they're functionally testing it every day. The cost is just the compute for a couple of VMs. It's not elegant, but it's a lever you can actually pull when the main console goes grey.



   
ReplyQuote
(@frankd)
Reputable Member
Joined: 2 months ago
Posts: 313
 

That "one throat to choke" line is a classic, but I always ask about the teeth behind it. What's the actual penalty in your SLA for a three-hour outage? Is it a meaningful service credit, or just a pro-rated refund for three hours of service?

For resilience, I haven't seen symmetrical active-active. The cost and complexity is prohibitive. The pattern I've helped implement is more about decoupling the access path from the security policy. We use one primary SASE vendor for the full suite, but we have a lightweight, always-on VPN configured with a second cloud provider *solely* for our incident management platform and a jump-box subnet. It's not for general traffic, it's for the team that needs to fix things when the main fabric breaks. The policy is so simple and static it never drifts.

It moves the goalpost from "keeping the business running" to "enabling the team that can get it running again." That's a much cheaper ask.


buyer beware, but buy smart


   
ReplyQuote
(@first_timer_evan)
Reputable Member
Joined: 4 months ago
Posts: 278
 

Yeah, that operational split is what worries me when I see these multi-vendor suggestions. It's one thing to have two consoles, but two entirely different skill sets? That's a huge ask.

So when you say it's the hidden labor cost, are you including the hiring and training to get there, or is it more about the constant friction between the teams trying to solve problems together? Feels like you'd need a dedicated integration role just to translate between them.

How did that retailer handle incidents that actually crossed both domains, like a remote worker needing access to a store system?



   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

The "one throat to choke" analogy always breaks down when you realize you're not doing the choking, you're just watching from the sidelines. Your point about the clean whiteboard architecture is exactly why these designs get approved, they look perfect in a diagram.

On your active-active question, I haven't seen a true, symmetrical setup that's manageable. The closest practical pattern I've observed is more about decoupling *access* from *policy*. One team runs their primary SASE vendor for everything, but they also maintain a dead-simple, always-on IPSec tunnel from their main offices to a secondary cloud (like GCP if their primary is AWS). It's not for user traffic, it's a dedicated conduit for data center/cloud management traffic.

This gives the operations team a guaranteed path to their control planes when the main SASE fabric is down. It's a tiny, static piece of config that never competes with the primary vendor's policies. It's not full resilience, but it gives you a lever to pull.



   
ReplyQuote