Skip to content
Notifications
Clear all

Showcase: My custom metrics for Cato tunnel jitter and loss in Grafana.

2 Posts
2 Users
0 Reactions
28 Views
(@data_pipeline_rookie_43)
Honorable Member
Joined: 5 months ago
Posts: 365
Topic starter   [#14591]

Hi everyone! I've been working on monitoring our Cato SDP tunnels more closely, and I wanted to share what I've built in Grafana for tracking jitter and packet loss. I'm pretty new to this whole data pipeline and monitoring game, so I'd love to get your thoughts and maybe some advice on how to improve it!

The thing is, the built-in Cato monitoring is great for a high-level view, but I wanted to create some custom dashboards that could show trends over time and maybe even alert us before users start complaining about call quality. I'm pulling the raw metrics from our Cato management console's APIs (which was a bit of a learning curve for me 😅) and then pushing them into a PostgreSQL database.

From there, I'm using Grafana to visualize. My main panel right now shows average jitter and loss percentage per tunnel over the last 24 hours, with a separate time series graph that lets me spot spikes. I'm also trying to calculate a simple "quality score" by combining those two metrics, but I'm not sure if my formula is giving us a useful signal.

Has anyone else here set up something similar? I'm especially curious about:
- How you're handling the data collection. Is a direct API pull the best way, or are you using some kind of log forwarding?
- Any pitfalls you ran into with the metrics themselves? I noticed the jitter values can be a bit jumpy.
- If you've set up alerts based on this data, what thresholds are you using?

Really excited to learn from the community here. I can share more details on the data flow if that's helpful!


rookie


   
Quote
(@jackr)
Trusted Member
Joined: 2 months ago
Posts: 31
 

Oh, this is a fantastic project to take on, and I totally get your desire to move beyond the built-in graphs. The Cato console is fine for a snapshot, but for real trend analysis and proactive alerts, you gotta roll your own.

I went down a very similar rabbit hole last year! I also found the API a bit fiddly to start with - especially getting consistent timestamps for the metrics. My first piece of unsolicited advice? Start logging the API response times and any errors alongside your tunnel metrics. It saved me when I realized my collection script was silently failing every few hours.

For the quality score, I'd be super curious to see your formula. I tried a few weighted combinations of jitter, loss, and latency, but what I found most useful was actually creating separate, simple thresholds for each. A single score often smoothed over the specific problem. Like, a latency spike might be fine for file transfers but murder for VoIP, you know?

Are you storing the raw sampled data points, or are you aggregating them (like, taking the mean) before they hit Postgres? That choice bit me later when I wanted to calculate percentiles.



   
ReplyQuote