Skip to content
Notifications
Clear all

Just built a system that pings Tunnel endpoints and pages us if they're down.

1 Posts
1 Users
0 Reactions
22 Views
(@auditor_abby)
Reputable Member
Joined: 6 months ago
Posts: 363
Topic starter   [#17827]

We rely heavily on Cloudflare Tunnels for internal application access. The problem is that Cloudflare's status page shows the overall service health, not the status of our specific Tunnel endpoints. An endpoint can fail silently due to a local network issue, a daemon crash, or a credential problem, and we wouldn't know until a user reports it. That's an unacceptable blind spot for a critical access path.

I built a simple monitoring system to close that gap. It pings the specific hostnames served by our Tunnels and triggers a PagerDuty alert if any become unreachable from the public internet. The core requirements were:
* It must monitor from outside our network (simulating a real user).
* It must not rely on the `cloudflared` daemon's own metrics, which might be up while the routing is broken.
* It must integrate with our existing incident response playbook.

The implementation is straightforward. A lightweight container runs a script that attempts HTTP/HTTPS connections to a list of our internal application FQDNs (e.g., `internal-app.corp.company.com`). It uses a simple curl-based health check with a strict timeout. Any non-2xx/3xx response or timeout triggers an alert. We run this on a separate, geographically distinct cloud VM we already use for external monitoring, ensuring we're testing the public DNS resolution and Cloudflare's routing, not our internal network.

Key configuration points we validated:
* The checks use the public DNS record, not an internal IP.
* The alert includes the specific failing hostname and the HTTP response code.
* We built in a short retry mechanism to prevent false positives from transient glitches before escalating to a page.

This isn't a replacement for full synthetic transaction monitoring, but it covers the base requirement: knowing if our Tunnels are fundamentally reachable. It also creates an audit trail of Tunnel uptime, which is useful for our internal availability reporting. The next step is to tie these alerts into our Cloudflare logs to correlate with any `cloudflared` error events.


Where is your SOC 2?


   
Quote