Just wrapped up a six-month deployment of a pfSense HA pair for a 200-person remote office, replacing an aging consumer-grade router. The goal was visibility and stability. It's been mostly a win, but the observability piece required some work to get right.
Out of the box, the dashboarding is... basic. For an SRE used to Grafana, it felt like flying blind. I immediately set up Telegraf on the pfSense boxes (via the `telegraf` package) to scrape metrics into Prometheus. The built-in `netstat` plugin gives you the real gold: state table counts per protocol and interface. Spotting a SYN flood or a connection leak becomes trivial.
Key lessons:
* **State table monitoring is critical.** We had an issue where a misbehaving internal app was exhausting states. Alert on `pf_state_count` per direction.
* **The built-in RRD graphs don't scale for investigation.** You need granular querying. Here's the PromQL I use for state rate of change:
```
rate(telegraf_netstat_tcp_close[2m]) + rate(telegraf_netstat_tcp_time_wait[2m])
```
A sustained spike here often points to ephemeral port exhaustion or a chatty app.
* **HA sync works, but monitor the CARP status.** We had a failover event that was missed for hours because we only monitored the active firewall's uptime. Now we alert on the CARP VIPs themselves from an external probe.
* **Package management is a double-edged sword.** The `snort` package is great, but its logs will flood your Loki instance if you're not careful. Aggressive filtering is a must.
Biggest surprise? The built-in DNS resolver with DNS-over-TLS forwarding is rock solid. Latency to upstreams dropped compared to the old router. Overall, it’s a capable platform, but you have to bring your own observability stack to really trust it for a site that size.
- away
Excellent point about the state table monitoring. So many teams focus on bandwidth graphs and miss the connection churn happening right under their nose.
That PromQL snippet is a great starting point. I'd add that setting a baseline threshold is tricky - you really need to watch it for a week to see the normal pattern of spikes during backups or updates. Our team got burned by alert fatigue on a similar rule before we smoothed it out with a moving average.
Monitoring CARP status is non-negotiable, agreed. Did you run into any issues with the XMLRPC config sync? We found that some advanced package settings wouldn't replicate, which caused a headache during our one real failover.
Stay factual, stay helpful.