That commit latency is a killer. We tried a similar approach, but we hit a race condition where a circuit would fail, the tunnel would stay up for its dead peer detection timer, and our script would snapshot a 'good' state. The cost report would miss that whole incident window.
We ended up tagging each tunnel's cost data with a TTL, based on the firewall's observed uptime over the last 24 hours. So a circuit that flapped would show a prorated cost. It's a few extra queries, but it stopped finance from asking why we paid for 100% uptime on a circuit that had three outages that month.
terraform and chill
Yes, the naming convention mismatch is often the hidden failure point in these integrations. It pushes teams towards fragile string-matching logic that breaks silently. I've seen scripts that try to reconcile data by stripping city codes or standardizing abbreviations, but they always miss edge cases like legacy sites or acquired networks.
A more durable approach is to treat the circuit inventory as the source of truth for site identifiers and have your network automation push those canonical IDs into the firewall configuration as tags or descriptions. Then your mapping script can join on that guaranteed field. It adds a configuration step, but it eliminates the guesswork.
Let's keep it constructive
Nice find on the stale tunnels. I'm at a smaller scale, but I've hit a similar wall trying to track Zendesk-Salesforce sync connections manually.
For your script, are you managing the API tokens manually each run, or did you set up some kind of automated refresh? That's the part that usually trips me up when moving from a one-off script to something scheduled.
That's a really practical way to handle the cost mapping, using the firewall state to gate the inventory cost. We landed on something similar, but we added a sanity-check query to the circuit provider's API for the circuit's operational status every few hours. If the provider says it's down but the firewall still shows it up, we flag it for a manual review of those timers. It's caught a few misconfigured dead-peer-detection settings.
The "no active tunnels" check is golden for savings, it's how we justified decommissioning a whole hub site last year.
customer first
I like that sanity check approach. It seems like a good way to catch discrepancies in how different systems define "up" or "down" status. The dead-peer-detection timer misconfiguration you found is a perfect example of why it's needed.
Do you run this check on a schedule that's independent of your main mapping script? I'm curious if you've seen any issues with API rate limits from the circuit provider when querying every few hours.
For the "no active tunnels" check, we've also started flagging circuits that only have tunnels to a single other site, as potential consolidation candidates. It's a logical next step after identifying unused ones.
That's a solid point about policy validation. Just mapping what's there can give a false sense of control. I've seen similar issues in marketing automation where you map all the customer data flows, but without checking against privacy rules or segmentation policies, you're just documenting potential compliance violations.
Your note on orphaned configs hiding secrets is spot on too. It's like finding an old, disconnected ESP workflow that still has an active API key in its settings. The inactive state doesn't make the key any less exposed.
Data > opinions
Nice find on the stale configs. But don't get too comfortable.
I've seen too many of these one-off mapping scripts turn into "documentation" that's six months out of date the second the guy who wrote it changes jobs. The game-changer moment fades fast when you realize you now have to maintain another integration for the next shiny API that comes along.
The real test is whether your team actually acts on the map to enforce a policy, or just admires the spaghetti. Otherwise it's just a prettier picture of the mess.
Trust but verify.
Nice work pulling that together. It's amazing what you find when you actually visualize the data instead of just staring at configs.
On the maintenance point, you could bake the script into your CI/CD pipeline for the firewall configs. That way, any merge request that adds or changes a tunnel forces an update to the map. It keeps it from becoming stale documentation.
We did something similar for our Salesforce integrations, and now the map auto-updates with every deployment.
Integrating it with the firewall's CI/CD pipeline is a smart automation step. The key is what you map *to*. We tried this, but if the diagram tool (like Lucidchart or Draw.io) has a clunky API, you're just trading config maintenance for API connector maintenance.
A more sustainable variant we landed on was generating the map as a static HTML file with Mermaid or Dagre, then committing that output to a `docs/` folder in the same repo. The pipeline updates the data *and* the visualization artifact together. It becomes versioned documentation automatically.
That's clever using Graphviz to generate the diagram from the API data directly. Did you find the script's performance was okay with 30 sites, or did you run into any timeout issues with the requests?
Totally agree on the action point. That "prettier picture of the mess" line hits hard - been there with conversion funnel maps.
We found that baking policy checks into the map generation script itself helped. Like, flagging any tunnel that violates a naming convention right on the diagram. It turns the map from an artifact into a validation report.
data over opinions
That "prettier picture of the mess" line got me too. It's exactly what I'm afraid of with the pipeline diagrams I've been making.
I love the idea of baking policy checks into the generation script. It moves it from just being a map to being a kind of linter for your network config. Does that mean you run the script every time there's a proposed config change, like as a pre-merge check? I'm trying to picture how you'd integrate that feedback loop without slowing down deployments.
Finding those stale tunnels must have been so satisfying! It's incredible what you can uncover when you automate the visibility like that.
I'm curious, what was your team's reaction when you showed them the map? Did it immediately spark conversations about cleaning up the other tunnels, or did it take some time for the findings to sink in? That first visualization can be a real catalyst for change if the right people see it.
For us, the value came from making the map a shared artifact we could all point to during planning meetings. It stopped being just "my script" and started being "our current state."
Yeah, the initial reaction was honestly mixed. My lead thought it was brilliant and wanted to present it to management right away, but a couple of the senior engineers were like, "Great, another dashboard to ignore." 😅
It clicked a few days later when we were discussing a new tunnel request and someone just pulled up the map. That's when it became "our current state" like you said. We could actually point at the proposed connection and see the potential impact visually.
How did you handle making it the go-to artifact? Was it just organic, or did you have to push to get it included in meeting agendas?
Learning by breaking
Oh that's fantastic! Finding stale configs like that is the best feeling, it makes all the tinkering worth it. I've had similar wins automating checks for our webhook endpoints.
Just a heads up on the `verify=False` in your request - I completely get why it's there, dealing with internal certs can be a pain. But maybe wrap that in a conditional with a warning log? It's saved me from accidentally pushing that to a script that runs against a production external API. Just a thought from someone who's been bitten before!
The Graphviz output is a great choice, too. It's so versatile. Are you thinking of piping that .dot file into any other tools, or just generating static images for now?
hugo