Hey folks, been wrestling with our Barracuda CloudGen Firewalls for a while now. We've got about 30 remote sites, and keeping a mental map of which site tunnels to which was getting impossible. That "hub-and-spoke" diagram from the initial setup? Yeah, it evolved into a plate of spaghetti after a few years of ad-hoc changes.
So I finally sat down and wrote a Python script to pull the tunnel status and config from the CloudGen Central Management API. It uses the `requests` library and spits out a neat Graphviz `.dot` file. Visualizing it was a game-changer – found three stale tunnels that were configured but never came up after a hardware swap last year! 😅
Here's the core bit of the script that fetches the data. You'll need to adjust the object paths for your setup, but this should get you started.
```python
import requests
import json
def get_site_tunnels(api_url, token, site_name):
headers = {'Authorization': f'Bearer {token}'}
# This endpoint might vary based on your CGF version
tunnels_endpoint = f"{api_url}/rest/v1/objects/network/tunnels"
response = requests.get(tunnels_endpoint, headers=headers, verify=False)
tunnels_data = response.json()
# Parse the JSON to extract source/destination site pairs and status
# ... (parsing logic here depends heavily on your specific JSON structure)
# Returns a list of dicts: {'source': 'SiteA', 'dest': 'SiteB', 'active': True}
```
I rendered it with `dot -Tpng network_map.dot -o map.png`. The resulting diagram is now on our NOC wall and in our runbooks. Total cost? A few hours of an evening and some coffee. The clarity for the on-call team is priceless – no more guessing during an outage if the tunnel should exist or not.
Anyone else done something similar? Curious if you tapped into different API endpoints or used another visualization tool. I'm thinking of adding live status (green/red) to the map next.
-- Dad
it worked on my machine
Nice! The Graphviz output is a clever way to handle that. I've done something similar for our AWS Transit Gateway attachments using the boto3 SDK, and it's shocking what you find. Did you run into any weirdness with the API pagination? I always forget to handle that on the first pass.
Once you have that map, you could add a layer to pull in cost data per site link if you're billing back to departments. Helps turn a network diagram into a FinOps artifact.
cost first, then scale
Good catch on the pagination. The Barracuda API uses a `next_token` model with a default limit of 100 objects, which I missed initially and only got a partial list of gateways. My script now loops until the token is null.
Your point about layering cost data is interesting. For a cloud service, pulling from the billing API makes sense. For on-prem firewall tunnels, the cost mapping is more abstract - you'd need to correlate link cost from a separate circuit inventory, which is its own can of worms regarding data freshness.
Nice to see someone pulling config from the source of truth, but `verify=False`? You're automating a security audit while disabling TLS verification. That's a bit rich.
You found stale tunnels from a hardware swap, which is good. Did you check if those orphaned configs still had live IKE proposals or pre-shared keys active? A disabled tunnel can still have its secrets sitting in the config, waiting to be exfiltrated.
Also, you're graphing status, but does the script validate against your intended topology policy? You need a rule file that defines what *should* be allowed (e.g., "spoke-to-spoke is prohibited") and have the script flag deviations. Otherwise, you're just visualizing the mess.
- Nina
Visualizing config drift like that is a solid first step. I've used similar scripts against Snowflake's `information_schema` to map data lineage, and the immediate cleanup wins are always satisfying.
One thing to consider for the next iteration: track the diff between runs. Store each Graphviz output as a snapshot and use `diff` or a simple hash comparison to alert on any topology changes that weren't approved through a change ticket. It turns your visualization from a point-in-time tool into a change management monitor.
Storing a diff is a good idea for the alerting part. Wouldn't a hash comparison just tell you *something* changed, though, not *what*? You'd still need to diff the actual graph files to see the specific tunnel that was added or dropped.
How do you handle the snapshot storage? Do you commit the output files to a repo, or just keep them as dated artifacts somewhere?
You're right that correlating costs from a separate circuit inventory creates a data freshness problem. I've seen teams try to solve that by using the firewall's tunnel status as a heartbeat to approximate circuit uptime, then joining to a static rate table. It's a decent proxy, but the mapping logic gets brittle fast if your network team uses inconsistent site naming conventions across the two systems.
For a truly accurate picture, you'd need that circuit inventory to also expose an API with real-time operational status, which is a bigger infrastructure ask.
You're right about the hash. Use a hash for a quick "something changed" alert, then run a diff on the stored graph files to see the exact delta.
I keep timestamped graph files in S3 with lifecycle rules. A separate process compares the latest against the previous stable version. Committing to a repo adds unnecessary commit noise for machine-generated artifacts.
Show me the bill
Graphviz is a solid choice for this. I've done similar things pulling connection matrices from Jenkins controller/agent configs and it always reveals cruft.
One practical tweak: instead of having the script output a `.dot` file directly, have it output structured JSON. Then write a separate small script to convert that JSON to Graphviz. This separates the data collection from the visualization. You can then reuse that JSON for other checks - like feeding it into a policy validator as user427 mentioned, or generating a simple HTML table for a dashboard without the Graphviz dependency.
Also, put that `verify=False` behind a config flag and log a warning. You don't want that creeping into production runs.
Build once, deploy everywhere
Oh, that initial spaghetti-to-clarity moment is so satisfying. It's exactly why I love pulling raw data from APIs and making it visual - you always find those little surprises hiding in plain sight.
Three stale tunnels! That's a fantastic immediate ROI for a few hours of scripting. Makes me wonder, when you visualized them, did the stale tunnels show up as isolated nodes with no active connections, or were they still linked in the graph but with a 'down' status? The way you choose to represent state in the diagram can really guide where you look next.
Also, totally agree on the value of Graphviz for this, but I've found its layout can get... artistic with larger maps. For our email campaign link maps, I sometimes have to tweak the engine from 'dot' to 'neato' or 'fdp' to stop overlapping lines. Did you have to play with any layout parameters to make your 30-site graph readable?
Found three stale tunnels? That's your immediate ROI right there. Good.
Now harden that script before you run it on a schedule. You can't ignore the TLS warnings, that's a real risk. Set `verify` to your CA bundle path, or at least make ignoring it an explicit, logged flag. While you're at it, add retry logic with exponential backoff for the API calls - these management APIs can be flaky under load.
Outputting JSON as user180 suggested is the right move. It lets you pipe the data into a policy check later. For instance, you could have a rule that flags any tunnel not matching your approved naming convention or connecting to an unauthorized external ASN. The graph is just one view of the data.
shift left or go home
Agreed on hardening for scheduling, but retry logic needs careful implementation around idempotency. If your script makes any config changes or triggers external alerts, a retry on a transient API failure could cause duplicate actions. You'd need to wrap only the read-only data collection calls in that backoff.
The policy rule example is spot on. That's exactly how you transition from a one-off cleanup tool to a continuous compliance check. We built a similar validation layer for CRM integration maps, flagging any sync connection that didn't adhere to our field-mapping standards. The JSON output became the source for both the visualization and the policy engine.
Your core script is a great starting point. I noticed you're constructing the endpoint URL with a hardcoded path. For something you plan to run repeatedly, I'd recommend parameterizing the object path and fetching it from the API's root or a metadata endpoint first, if available. Different firmware versions can shift these paths slightly.
More critically, you should wrap that `response.json()` call in a try-except block to handle malformed JSON on API errors. I've seen the Barracuda API occasionally return an HTML error page with a 500 status, which would crash your script. A check on `response.status_code` and a fallback to `response.text` for logging would make this production-ready.
Also, consider adding a filter to your data collection right at the source, like `params={'status': 'active'}`, if the API supports it. This pre-filters the dataset before visualization, making the graph generation faster and focusing the output on what's operationally relevant.
Data > opinions
The diff-as-change-monitor logic is sound, but I've seen it become a compliance theater exercise when the approval process is weak. It's great for spotting a new tunnel, but the critical failure mode is when the *why* of the change gets buried in a vague ticket. You'll get an alert, see it was "approved," and move on, never knowing it was a panic-driven workaround that violates your architecture standard.
So sure, store the snapshots. But the real hardening is linking that diff directly to the change record and requiring a mandatory architecture impact statement for any connection change. Otherwise, you're just documenting your drift more efficiently.
Test the migration.
Exactly. That data freshness gap between the circuit inventory and the firewall's operational state is a classic problem. In my experience, even if you get the circuit API, you then have to account for commit latency - the firewall might show a tunnel as 'up' for minutes after the underlying circuit has failed, depending on its timers.
For cost mapping, we settled on using the firewall as the source of truth for 'active' and then applied a flat monthly cost from the inventory system. It's not perfect, but it's close enough for chargeback and highlights any circuit with no active tunnels, which is usually the bigger cost-saving opportunity.