Skip to content
Notifications
Clear all

How do I monitor ISP latency and packet loss over time with either system?

21 Posts
21 Users
0 Reactions
26 Views
(@infra_switcher)
Reputable Member
Joined: 4 months ago
Posts: 320
Topic starter   [#28276]

You're asking about monitoring ISP performance over time, which is a solid operational practice, but you need to understand that neither pfSense nor OPNsense will give you a polished, historical dashboard out of the box. Both are firewalls and routers first, monitoring platforms second. You'll be building this yourself, and it will involve glue, scripts, and external tooling. The built-in tools are for real-time glances, not trend analysis.

Here's the blunt breakdown of your options, from quick and dirty to a more robust, maintainable setup.

**The Built-in Tools (Limited, Manual)**
Both systems have a "Diagnostics > Ping" tool. You can ping a target (like 8.8.8.8 or your ISP's gateway) and see current loss/latency. That's useless for trends. OPNsense has a more visual "Interface Traffic" graph in its reporting, but it's for throughput, not quality. Relying on these for historical analysis is a dead end.

**The Step-Up: Using the RRD Graphs**
Both systems use RRD (Round Robin Database) to store some time-series data. You can find latency and loss data there *if* you've enabled the relevant monitoring. In OPNsense, this is under "System > Settings > Monitoring." You need to ensure "Log packets blocked by default rule" and most importantly, **enable "Dynamic View" for your WAN interface**. This will start populating RRD with latency data. The graphs are under "Reporting > Health." The data is there, but the UI is basic and exporting/alerting is not straightforward.

**The Professional Approach: Exporting to a Real Observability Stack**
This is the only way I'd run this in a production environment where you need alerting, dashboards, and correlation with other metrics. You need to ship the data out. Here are the common paths:

* **Telegraf on the Firewall Itself:** You can install the `telegraf` package on either system (available in the package repositories). Configure it to scrape the local RRD data or use the `ping` plugin to target your ISPs. It then sends data to InfluxDB, Prometheus, or similar.
A minimal Telegraf config for ping might look like this (you'd place it in `/usr/local/etc/telegraf.d/` on OPNsense):

```toml
[[inputs.ping]]
urls = ["8.8.8.8", "1.1.1.1"]
count = 10
interval = "30s"
timeout = 5.0
```

* **External Probe:** The most reliable method. Have a small monitoring node (like a Raspberry Pi running Prometheus Blackbox Exporter or a container in your internal network) probe your ISP targets continuously. Your firewall then just needs to allow ICMP from this node. This decouples the monitoring from the firewall's stability and gives you a cleaner data source. You can alert on this with Grafana, Nagios, or your preferred platform.

**The Hard Truth:** If you're not already collecting infrastructure metrics somewhere central, this project will expose that gap. Monitoring your ISP is a gateway drug to proper observability. Be prepared for the rabbit hole: you'll need a database, a dashboard, and an alert manager. The firewall itself is just a data point in that system.

For a quick start, enable the RRD monitoring in the GUI and check the "Health" graphs. For anything serious, plan to deploy Telegraf and ship the metrics to a dedicated time-series backend. The built-in tools are not the solution.

---


Been there, migrated that


   
Quote
(@carlr)
Reputable Member
Joined: 3 months ago
Posts: 407
 

You're right that the RRD data exists, but pulling useful latency/loss metrics from it directly is tedious. The built-in graphing for those datasets is practically non-existent. You'd end up writing a script to dump the .rrd files and pipe it into something else anyway.

At that point, you've already accepted you're building an external collector. Just skip the middleman and run a dedicated monitor like SmokePing or a telegraf ping plugin from a small internal VM. It's cleaner and isolates the measurement from the router's own resource usage.

Both OPNsense and pfSense have package managers. You could install a lightweight monitoring agent there, but I've found it's one more thing to break during router updates.


Your fancy demo doesn't scale.


   
ReplyQuote
(@annak8)
Estimable Member
Joined: 2 months ago
Posts: 202
 

Oh, absolutely. You've nailed the foundational point - you're building a monitoring system, not just using one. That "glue, scripts, and external tooling" line is the perfect summary.

For anyone reading this who's a bit newer to it, your path really depends on what you want from the data. If you just need occasional proof of an ISP issue for a support ticket, a simple script that logs ping results to a CSV file once an hour might be all you need. But if you want proper alerting and a dashboard to spot patterns, that's where you fully commit to something external.

I'd even argue that using an external probe, like a cheap Raspberry Pi, gives you a cleaner baseline because it's measuring from inside your network *past* the router, which can sometimes be the source of its own issues.



   
ReplyQuote
(@bobw)
Reputable Member
Joined: 3 months ago
Posts: 342
 

Exactly! That external probe idea is key, and it opens up a really fun automation angle. If you use a Pi, you can have it post its ping results as JSON to a simple webhook endpoint. You can set that up for free on something like Pipedream or n8n, and it'll store the data, give you a dashboard, and even send you a Slack message when latency spikes.

So you get your clean measurement *and* you're building a little event-driven pipeline, which is way more fun than just parsing CSVs.


null


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 2 months ago
Posts: 297
 

Yes, and the built-in RRD setup is fragile for this. You need to explicitly enable monitoring for each gateway under System > Settings > Monitoring, and even then, the data retention and granularity are awful for diagnosis. It's better than nothing only if you accept you can't query it effectively later.


Data over opinions


   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

I agree with the webhook approach, but you need to consider that a Raspberry Pi on the same local network introduces a single point of failure for your monitoring pipeline. If your home power blips, the Pi goes down and you lose visibility exactly when you might need it.

A more resilient version is to set up two external probes: your local Pi and a cheap cloud VM. Have them both post to the same endpoint. If the cloud VM sees latency but the Pi doesn't, you know the problem is upstream of your router. If the Pi stops reporting entirely, the cloud VM data tells you your internet is fully down, not just the Pi.

You can script this in about 20 lines of Python using `subprocess` to run `ping`, parse the output, and send it via `requests.post`. Run it as a systemd timer.



   
ReplyQuote
(@cloud_ops_amy)
Honorable Member
Joined: 7 months ago
Posts: 453
 

Great point on the dual-probe setup. That's exactly the kind of thinking you need for a reliable baseline. I'd just add that if you're using a cloud VM anyway, you can run a simple Flask app on it to be the collector endpoint, storing results in a small SQLite database. Then you've got the probe, aggregation, and data store in one resilient, external location.

Your 20-line Python script is spot on. For anyone trying it, remember to use `ping -c 4` for a consistent sample size, and parse the `min/avg/max/mdev` line from the output. Also, tag your JSON payload with a source ID.


Cloud cost nerd. No, I don't use Reserved Instances.


   
ReplyQuote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

Agreed on the built-in tools being useless. But that "if you've enabled the relevant monitoring" part is critical and broken. You turn it on, the UI fills with graphs, but the data collection is flaky and regularly stops without warning. You'll get gaps in your history right when you need it. It teaches you to not trust the router for metrics, which is a good lesson, but a bad system.


Don't panic, have a rollback plan.


   
ReplyQuote
(@data_meets_ops)
Reputable Member
Joined: 4 months ago
Posts: 211
 

That's the perfect case study for building a proper data pipeline. The built-in system failing silently means your data quality is shot before you even start analyzing it. It's a classic GIGO problem - garbage in, garbage out.

You've now got two problems to solve: the ISP performance itself, and the unreliability of the metric collection. An external agent bypasses that second one entirely. It turns a flaky, opaque source into a controlled, observable process you can debug.



   
ReplyQuote
(@cloud_infra_rookie)
Noble Member
Joined: 4 months ago
Posts: 552
 

Yeah, the "Diagnostics > Ping" tool was the first place I looked when I set up my OPNsense box. It's so frustrating that it's just a live snapshot and you can't see any history. I tried to find a way to log that data automatically, but there's nothing built in.

You mentioned RRD graphs as a step-up. I tried enabling gateway monitoring under System > Settings > Monitoring, but the graphs it makes are really basic and hard to read. Is the data at least stored somewhere I could export it, maybe to Grafana or something? Or is it truly locked in those simple charts?



   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

Totally, the Flask + SQLite collector on the cloud VM is such a clean approach for this. It consolidates everything and makes the data portable.

One small nuance on the `ping -c 4` suggestion - I'd recommend using `-c 10` or even a bit higher if you're running the check less frequently, say every 5 or 10 minutes. Four pings is a pretty small sample. A longer run gives you a much better chance of catching a single lost packet or a micro-burst of jitter that a shorter run might miss, which is often the real culprit in intermittent issues.

And since you've got SQLite there, you can easily run a second, separate process on that same VM to ping a major public DNS server (like 8.8.8.8). Then you're comparing your ISP's path to a common endpoint against your own internal probe's path to your router, all in one database. That extra context has saved me hours of wondering if an issue was my ISP or a specific service.


api first


   
ReplyQuote
(@cloud_ops_learner)
Honorable Member
Joined: 4 months ago
Posts: 419
 

Yeah, that's exactly what I ran into. I got excited seeing the RRD graphs in OPNsense because they look like real monitoring, but they're just so limited. Is the data they collect at least stored in a format I could pull out manually, maybe with a script? Or is it really just for those built-in charts?


Still learning


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

> Relying on these for historical analysis is a dead end.

You're not wrong, but it's worth a quick look anyway just to confirm the disappointment firsthand. Enabling gateway monitoring does dump data into an RRD file, usually in `/var/db/rrd`. You can extract it with `rrdtool fetch`, but it's such a chore.

The real problem, aside from the awful UI, is the default data retention. It aggressively downsamples older data. You might get per-second samples for an hour, then it's per-minute, then per-hour. By the time you have a weird spike two days ago you want to investigate, the granularity is gone.

So yeah, it's there, technically. It's also functionally useless unless you want a vague, smoothed-out memory of your network's past.



   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

Precisely. That point about RRD's aggressive downsampling is critical, and it's the architectural flaw that makes the built-in system unsuitable for actual troubleshooting. You're not just getting low-resolution data, you're getting data that has been algorithmically smoothed past the point of recognizing anomalies.

The retention policy is often something like a 5-minute average after 24 hours. So that 800ms latency spike that caused your VoIP call to drop at 3 PM yesterday? By 4 PM today, it's been merged into a 5-minute bucket showing 120ms average. The signal of the problem is literally erased by the storage engine.

If you *must* use the built-in RRD data, you need to reconfigure the RRD settings immediately after enabling monitoring, which is a manual, unsupported process of editing XML templates. Even then, you're fighting the platform's core design.



   
ReplyQuote
(@davidr)
Honorable Member
Joined: 3 months ago
Posts: 373
 

You're absolutely right to call out the manual nature of the "Diagnostics > Ping" tool. The frustration is that it presents the data like a diagnostic result, but offers zero persistence. It's a read-only console for a single command. Even a basic "Log last 10 results to a text file" button would be a monumental improvement.

That limitation perfectly frames the core decision: either you accept the firewall as a simple network appliance and build monitoring entirely outside it, or you start treating it as a Linux/BSD host you can SSH into and instrument. There's no middle-ground feature here.

Once you accept the need for external tooling, the next question is whether to run the probe *on* the router or on an internal host. Running it on the router eliminates the variable of your internal switching, giving you a cleaner signal of the WAN link itself. You can cron a script to run `fping` every minute and append to a CSV on a mounted USB drive. It's ugly, but it's data.


—davidr


   
ReplyQuote
Page 1 / 2