While my primary focus is cloud infrastructure, our recent deployment of a Sophos XGS for a hybrid office presented a familiar problem: visibility into operational costs. In this case, the "cost" was not dollars, but reliability minutes lost during WAN failover events. The built-in reports showed the failover occurred, but provided little granularity on duration, frequency, or pattern.
I built a simple dashboard to treat WAN link stability as a measurable resource. The goal was to move from "the firewall failed over last night" to concrete data for capacity planning.
The methodology was straightforward:
* Configured the XGS to send syslog data to a central collector.
* Parsed logs for specific event IDs related to WAN interface state changes (link up/down).
* Used a time-series database (Prometheus) to count and time the events.
* Visualized the data in Grafana.
The resulting dashboard provides immediate insight into:
* Total failover events per WAN link over configurable timeframes.
* Mean time to restore (MTTR) for primary link recovery.
* Patterns that might indicate chronic ISP issues at certain times.
This approach allows for data-driven conversations with ISPs about SLAs and informs decisions on whether a redundant internet connection is a necessary expenditure or an underutilized asset. The principle is identical to rightsizing a cloud instance: you cannot optimize what you do not measure.
Optimize or die.
CloudCostHawk