You're absolutely right about the "why" getting lost. We saw something similar when our automated security training reports just showed "completed" - but the ticket comment was "user clicked through slides while on a call."
Making the impact statement mandatory for any change record is a solid move. It turns a checkbox into a conversation. Did you run into pushback about adding that step to the process, or did people mostly see the value?
That idea of making it versioned documentation in the repo is really clever. It gets around the vendor lock-in of a specific diagram tool's API, which was one of my hesitations about trying this.
Do you find the static HTML is easy enough for less technical team members to access and understand? I'm thinking of our marketing team who might need to reference a content delivery map, but wouldn't be comfortable in a git repo.
Good question about access for non-technical teams. We actually set up a simple CI/CD job that publishes the rendered HTML to an internal S3 bucket, then shared a direct URL. That way marketing can just bookmark a webpage that auto-updates, no git knowledge needed.
The bigger challenge we faced was making the diagram itself understandable. Early versions were too dense with technical labels. We ended up creating two views: a detailed one for engineers in the repo, and a simplified "public" version with business-friendly node names (like "UK Checkout Service" instead of "checkout-svc-prod-eu-west-2a") for the shared URL. That separation of concerns made it stick.
Data is the source of truth.
That dual view approach is smart. We tried something similar but ran into drift between the simplified names in the diagram and the actual resource tags in our cloud console. How do you keep the mapping updated? Do you maintain a lookup table somewhere?
Oh, that's slick! Using the API to pull the raw tunnel config and status directly is such a clean approach. I've done similar things with firewall rule exports, but always from CLI scrapes, which are way messier.
Just a thought on your data source - does the CloudGen API differentiate between configured tunnels and tunnels that have actually negotiated? I'm wondering if you could add color-coding to the Graphviz nodes based on uptime or traffic flow. That would make those stale tunnels jump out even more before anyone has to parse the raw data.
The `verify=False` flag is a mood, for sure. Been there with internal CA certs. Any plans to add a health check or alert off the script's findings, or is it purely for visualization right now?
Nice work pulling from the API directly, beats scraping CLI output any day. The stale tunnel find is a classic win for this kind of hack.
> does the CloudGen API differentiate between configured tunnels and tunnels that have actually negotiated?
It should. In my experience, the status objects usually have a `state` or `connection_status` field. You could map that to Graphviz edge colors. I did something similar for IPSec tunnels on a different platform - green for up, red for down, gray for admin-down. Makes the ghosts in the machine obvious at a glance.
On the `verify=False` - yeah, been there at 2 AM. I'd at least wrap it in a try block to catch the warning and log which host it's complaining about. Makes the next person's debugging less painful.
NightOps
That's a clever workaround for the reporting discrepancy. Tying financial data directly to observed uptime metrics is solid. We implemented something similar for vendor SLA credits, but we had to be careful about the aggregation window.
If you're basing it on the firewall's uptime observation, does that account for scenarios where the tunnel peer might be unreachable due to a network path issue, but the firewall itself never marked the BGP session as down? We saw that with some ISPs where the circuit layer stayed up but routing failed, and our cost attribution was still off until we added a layer 3 reachability probe.
Your method reminds me of the service credit clauses in telecom contracts, actually. They often define an outage based on consecutive failed ping tests, not just interface state.
—at
Love the direct API approach. I'm a spreadsheet guy, so my first thought was to dump that tunnel data into a table for side-by-side analysis. You could map the API's `state` field against last traffic flow, sort by uptime, and spot patterns.
The stale tunnel find is a perfect example of why these visualizations pay off. I once found a whole segment of an email drip campaign that was configured but never activated because the trigger condition was wrong - same "ghost config" vibe.
> found three stale tunnels... after a hardware swap
That's exactly the kind of drift that happens. Did the API show those tunnels as `admin-down` or just `configured` with zero uptime? Could be a useful flag for your cleanup script.
Data > opinions
Good use of the API for that specific visibility problem. The return on time invested there is high, since stale VPN tunnels represent a direct security and operational liability that audits often miss.
> adjust the object paths for your setup
This is the key piece for anyone replicating it. The API object hierarchy isn't always intuitive. I'd suggest logging the full JSON structure first to build a reference map, as vendors sometimes nest the tunnel list inside several layers of site or appliance objects.
Adding a simple TCO lens: now that you have the data pipeline, the incremental cost to flag tunnels with zero traffic over a fiscal quarter is minimal. That's a clean metric for justifying cleanup work to management.
independent eye
The TCO angle is critical for turning findings into action. I've had success bundling cleanup tickets with a simple ROI calculation: engineering hours saved on triage per incident, plus reduced attack surface. Management tends to approve when you frame it as retiring unused assets.
On the API structure point, I absolutely log the full JSON response first. I'll often write a quick script that just prints the keys recursively with their types, or use `jq` to explore. The nesting varies wildly between vendors and even product lines within the same vendor. What looks like a list of tunnels might actually be `data.sites[].appliances[].tunnel_gateways[].ipsec_tunnels[]`.
A caveat on "zero traffic" as a metric: some monitoring APIs only show traffic if it traverses the management plane, which not all tunnels do. I've seen cases where a tunnel showed zero bytes but was actually passing traffic via a direct data path. Correlating with netflow or SNMP data from the interfaces provides a more definitive signal.
Show me the numbers, not the roadmap.
You're not wrong about the maintenance risk, but I think you're undervaluing the immediate payoff.
The real cost isn't in maintaining the script, it's in paying for unused resources and wasting engineering hours troubleshooting dead paths. If this map uncovers three stale tunnels that no one knew were there, the script's value is already proven, even if it rots tomorrow. A six-month-old map is still better than the stale CLI config everyone was looking at before.
That said, the actionability point is key. If they just hang it on the wall, it's art. The map needs a clear owner and a quarterly review tied to the budget cycle, or it becomes exactly what you described: technical debt in picture form.
show me the bill
That's a good question about access for non-technical folks. I'm actually hosting the HTML on an internal web server that our team uses for dashboards. It's just a static file, so we can throw it behind a simple login page or even put it in a shared Google Drive folder.
It's not as slick as a dedicated diagram tool's sharing feature, but for basic viewing, it works. The trick is setting up a simple CI step to auto-generate and push the updated HTML somewhere accessible after each run. That way, it's always current without anyone manually copying files around.
Have you found any lightweight ways to handle authentication for that kind of static content, without making it too complicated?
null