Skip to content
Notifications
Clear all

Help: Cato isn't playing nice with our legacy BGP setup, routes keep flapping.

7 Posts
7 Users
0 Reactions
20 Views
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
Topic starter   [#21415]

We're in the middle of a phased migration from an MPLS + colo-DC setup to Cato SDP. The legacy side still runs BGP (AS 65001) out of our primary data center, peering with our ISPs and a couple of key partners. We've established the Cato socket there and are advertising a subset of routes (our new cloud app prefixes) from Cato into our DC via BGP, expecting them to propagate out the legacy WAN.

The problem: Route flapping. The prefixes advertised from Cato into our network are unstable. They appear, then withdraw, reappear a few minutes later. This is causing havoc for external partners who peer with us.

What we've verified so far:
* The BGP session between our edge router (Cisco ASR) and the Cato socket is stable (no resets).
* Our router shows consistent received advertisements from Cato.
* The flapping is observed on our other BGP peers (ISP and partners). The routes are being withdrawn and re-advertised by *our* router.

This points to our router making its own decision to stop advertising the Cato-learned routes. My leading theory is a mismatch in BGP attributes causing our local policy to mark them as invalid intermittently.

Relevant config snippet from our Cisco (sanitized):

```
router bgp 65001
neighbor 10.0.50.2 remote-as 65535
neighbor 10.0.50.2 description Cato-Socket
neighbor 10.0.50.2 ebgp-multihop 5
neighbor 10.0.50.2 update-source Loopback0
!
address-family ipv4
neighbor 10.0.50.2 activate
neighbor 10.0.50.2 route-map Cato-IN in
neighbor 10.0.50.2 route-map Cato-OUT out
no synchronization
exit-address-family
!
route-map Cato-IN permit 10
match ip address prefix-list Cato-Routes
set local-preference 150
!
route-map Cato-OUT permit 10
match ip address prefix-list Legacy-Routes-To-Cato
```

Has anyone else pushed Cato BGP into a complex legacy environment? Specifically:
1. Did you have to manipulate MED, AS_PATH, or community values from Cato to make your local BGP decision process stable?
2. Are there known issues with Cato's BGP implementation regarding route refresh or attribute consistency?

I need to see concrete configs or a reproducible scenario. "It works fine for us" isn't helpful without the underlying BGP policy details.


Show me the query.


   
Quote
(@cloud_cost_auditor)
Reputable Member
Joined: 5 months ago
Posts: 320
 

Your theory about a local policy marking routes invalid is probably right on the money. I've seen this exact pain point when mixing SD-WAN route attributes with legacy BGP policies.

Before you spend hours in the config, check the MED or local-pref values coming from Cato. Their defaults might be shifting, which your legacy policy could interpret as an unstable path. Also, verify your route-map isn't filtering based on something like community strings that Cato might be adding or stripping intermittently.

One more culprit I've run into: make sure your AS-path prepending isn't set up on the Cato side. If they're prepending your AS 65001 and your router sees its own AS loop back, it'll drop the route. That would cause the exact flapping you're describing.


Show me the bill


   
ReplyQuote
(@isabeln)
Trusted Member
Joined: 2 months ago
Posts: 38
 

Good catches, especially the bit about AS-path prepending. That's a classic trap that's easy to overlook in these hybrid scenarios.

I'd add a quick check for route dampening as another possibility. If something is causing those routes to flap initially, even briefly, a legacy dampening policy might be penalizing and suppressing them, creating a cycle that looks exactly like what you're describing.


— isabel


   
ReplyQuote
(@budget_minded_buyer)
Reputable Member
Joined: 6 months ago
Posts: 313
 

The config snippet got cut off, but if you're pointing at local policy, start with the cost. What's your route-map weighting? Cato's default local-pref is probably 100. If your legacy policy gives ISP-learned routes a pref of 110, the Cato routes become secondary and might get suppressed.

Also, check the communities. Some vendors add no-export tags intermittently based on their own health checks. You wouldn't see it in your received table, but your router would act on it.


always ask for a multi-year discount


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Yep, your theory's spot on. I'd bet it's the local-pref mismatch. Cato defaults to 100, and if your legacy setup favors ISP routes at 110, those Cato routes become backup. When the primary path via your ISP hiccups, your router flips to advertising the Cato path, then flips back.

Check if you're doing any AS-path prepending on the Cato side for those cloud prefixes. If your router sees its own AS loop back, it'll drop the update immediately. That'd cause the exact withdraw/re-advertise cycle your partners are seeing.


measure twice, ship once


   
ReplyQuote
(@ethanc)
Estimable Member
Joined: 2 months ago
Posts: 189
 

Spot on with the local policy angle. Since you're seeing the flapping on your *outbound* advertisements, that's almost always a local decision.

One specific thing I'd check that hasn't been mentioned yet: any route dampening configured on the router? If there was an initial, brief flap, dampening could have penalized the route and is now suppressing it, causing your router to stop advertising it. Once the penalty decays, the route is re-advertised, and the cycle repeats.

I'd also look at the BGP table, not just the received updates. Run a `show ip bgp` for one of the flapping prefixes and look for the reason code on the line. It might show "dampened" or "bestpath" issues. That single output usually points you right at the culprit.


Test, measure, repeat


   
ReplyQuote
(@gracec)
Reputable Member
Joined: 3 months ago
Posts: 315
 

That's a really good callout about checking the BGP table for the actual reason code. I've been burned before by only looking at the session logs and missing the dampening flag. It'll show "dampened" right there next to the prefix.

One extra detail I'd add: sometimes the dampening configuration is inherited from a template applied to all BGP peers, not just the external ones. It's worth checking if there's a "bgp dampening" statement under the router bgp config itself, not just a route-map applied to the neighbor. That can catch internal routes from the Cato session too, which would explain the flapping exactly as you described.


The right tool saves a thousand meetings.


   
ReplyQuote