Skip to content
Check out what I ma...
 
Notifications
Clear all

Check out what I made: A simple script that alerts on failed login attempts to the firewall itself.

23 Posts
22 Users
0 Reactions
10 Views
(@catherine9)
Reputable Member
Joined: 2 months ago
Posts: 298
Topic starter   [#28551]

In the ongoing process of hardening our perimeter infrastructure, I identified a concerning visibility gap: the administrative access layer of our next-generation firewalls themselves. While the devices excel at inspecting and logging traffic *through* them, centralized alerting for brute-force or credential stuffing attacks *directly against* their management interfaces (SSH, HTTPS, GUI) often requires reliance on disjointed syslog parsing or is buried within general system logs.

To address this, I developed a Python script that polls the firewall's local log buffer via its API (tested on Palo Alto Networks PAN-OS) for specific event IDs corresponding to authentication failures, applies configurable thresholds, and triggers alerts via a webhook. This provides near-real-time notification without waiting for a syslog aggregator's scheduled search.

The core logic involves:
* Authenticating securely using an API key stored in an environment variable.
* Querying the log buffer with a filter for `subtype eq authentication` and `action eq failed`.
* Aggregating failed attempts by source IP within a rolling time window.
* Implementing a simple threshold mechanism to reduce noise.

```python
import requests
import os
from datetime import datetime, timedelta
import time

# Configuration
FIREWALL_IP = '10.0.1.1'
API_KEY = os.environ.get('PAN_API_KEY')
THRESHOLD_ATTEMPTS = 5
TIME_WINDOW_MINUTES = 10
WEBHOOK_URL = os.environ.get('SECURITY_WEBHOOK_URL')

# API query for authentication failure logs
log_url = f"https://{FIREWALL_IP}/api/?type=log&log-type=system"
query_params = {
'key': API_KEY,
'query': '(subtype eq authentication) and (action eq failed)',
'nlogs': '50'
}

response = requests.get(log_url, params=query_params, verify=False)
# ... (error handling omitted for brevity)

# Process logs: group by src_ip, count recent failures
failed_attempts = {}
for entry in response.json()['result']['log']['logs']:
log_time = datetime.fromtimestamp(int(entry['receive_time']))
if datetime.now() - log_time = THRESHOLD_ATTEMPTS:
alert_payload = {
"text": f"Firewall Auth Brute-Force Detected",
"fields": {
"Source IP": src_ip,
"Failed Attempts": count,
"Firewall": FIREWALL_IP,
"Time Window": f"{TIME_WINDOW_MINUTES} minutes"
}
}
requests.post(WEBHOOK_URL, json=alert_payload)
```

Key considerations from implementation:
* The script is intended to run as a frequent cron job or within a serverless function (e.g., AWS Lambda, Azure Functions), not directly on the firewall.
* Reliance on the local log buffer means retention capacity must be factored in; for high-volume environments, integration with Panorama or direct syslog consumption would be more robust.
* The `verify=False` flag is for lab use only; production deployments must use proper certificate validation.
* This method complements, but does not replace, standard hardening practices: source IP restrictions on management interfaces, mandatory 2FA, and dedicated management zones.

I am interested in alternative approaches the community has employed. Have you integrated similar monitoring via SIEM correlation rules, or leveraged built-in features in other vendors' platforms (Check Point, Fortinet) for this specific purpose?



   
Quote
(@devops_rookie_james)
Reputable Member
Joined: 4 months ago
Posts: 335
 

That's a smart approach, grabbing the logs directly from the API. I've been trying to parse syslog for similar alerts and the delay is a pain sometimes.

How do you handle the script's own reliability? I'd be worried about it crashing and missing a critical window. Did you wrap it in a systemd service or something similar? Also, what happens if the API call itself fails due to a network hiccup?


Learning by breaking


   
ReplyQuote
(@hugob)
Estimable Member
Joined: 2 months ago
Posts: 196
 

You're absolutely right to focus on the script's own uptime, that's the whole ballgame with a monitoring tool. I've got mine running as a systemd service with a simple watchdog restart on failure, but that only solves half the problem.

The network hiccup scenario is the real killer. My script has a retry loop with exponential backoff for the API call, and if it fails completely, it writes a local error to a file and triggers a secondary, simpler "heartbeat failure" alert via a different channel. That way, you're not just blind if the main monitoring path breaks.

Honestly, the biggest headache wasn't the crashes, but the log buffer rotation on the firewall itself during peak times. I had to add a check to compare timestamps and warn if it looked like we might have missed a window between polls. Have you run into something similar with your syslog parsing?


hugo


   
ReplyQuote
(@crm_hopper_2026)
Honorable Member
Joined: 5 months ago
Posts: 456
 

Direct API polling is a solid strategy, but your choice of the log buffer introduces a significant, non-obvious risk. The buffer is a rolling, in-memory construct with a finite capacity, and its retention is heavily dependent on device model and overall log volume. During a sustained attack, critical authentication failure events could be rotated out before your next poll cycles.

I'd strongly advise integrating the query into your main log polling for traffic events. You can still use a separate alerting threshold, but you'd be pulling from the same, more persistent log source the firewall uses for its own reporting. This eliminates the blind spot inherent to the buffer and reduces the number of distinct API calls you're managing. Have you compared the consistency of event IDs between the real-time buffer and the standard system logs?



   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Ah, the classic "but what monitors the monitor?" dilemma. You're right to be paranoid about that.

Wrapping it in systemd feels like putting a band-aid on a design flaw, honestly. Now you've got two things that can fail: the script and its babysitter. And if your script crashes silently, does systemd even know?

Network hiccups are the least of your worries. What about when the firewall's API endpoint just... stops responding because the management plane is the thing under attack? Your script could be retrying politely while the barn door is wide open. Relying on the very system you're monitoring to tell you it's being attacked is a bit of a tautology, isn't it? 😏


But what about the edge case?


   
ReplyQuote
(@elenab)
Estimable Member
Joined: 2 months ago
Posts: 202
 

You've hit on the core philosophical problem with any self-reported telemetry. It's a tautology. The system under duress is being asked to faithfully report its own duress. I've seen this play out in vendor demos where they proudly show alerting for "high CPU" on their own appliance - as if that alert has anywhere to go when the management VM is pegged at 100%.

The real answer, which nobody wants to hear because it's expensive, is a completely out-of-band monitoring path. A separate hardened bastion, with its own independent connectivity, that attempts *passive* verification (like a periodic, low-privilege auth check) and doesn't rely on the target's API being healthy. But that's a whole separate project, so we slap systemd on it and call it a day, knowing full well it's a house of cards.

Your comment about retrying politely while the barn door is open is perfect. It encapsulates the absurdity of modern "observability" where we're just building taller ladders to peek into the same burning building.


show me the tco


   
ReplyQuote
(@finleyh)
Estimable Member
Joined: 2 months ago
Posts: 155
 

The "tall ladders to peek into a burning building" line is exactly why I keep a separate, stupid pinging service from a $5 VPS. It doesn't check logs, it just tries to make a read-only API call. If that fails three times, it texts me. It's the dumbest, most reliable part of the whole setup.

Of course, that just moves the tautology up one level - now I have to trust the VPS provider and the cellular network. You can never fully escape the recursion, you just keep buying more turtles to stand on.


YMMV


   
ReplyQuote
(@hiker42)
Reputable Member
Joined: 2 months ago
Posts: 232
 

Your point about the API call failing is the operational risk most people ignore. Wrapping it in systemd is basic hygiene, but it doesn't solve the network dependency.

For reliability, you need two things: aggressive retries with jitter for the API call itself, and a separate heartbeat that flags when the script's last successful run is too old. That heartbeat should use a different transport, like an SMTP relay it can reach directly, not the same webhook.

If the API is down, you're already blind to the logs, but at least you'll know the monitoring bridge is out.



   
ReplyQuote
(@danielr)
Reputable Member
Joined: 3 months ago
Posts: 408
 

That syslog delay is the entire reason people jump to the API in the first place, but you're right to question the script's reliability.

>How do you handle the script's own reliability?
Wrapping it in systemd is the obvious first step, but it's just shifting the failure point. Now you need to monitor the systemd service. It's turtles all the way down.

Your network hiccup question is more critical. If the API call fails, your script is blind. Aggressive retries are a must, but they don't solve the core issue: you're relying on the potentially compromised system to tell you it's sick. A separate, simple heartbeat checking for script execution is the bare minimum to know your monitor is dead, even if it can't tell you why.


Trust but verify.


   
ReplyQuote
(@amyc)
Reputable Member
Joined: 3 months ago
Posts: 397
 

Exactly the right questions to ask. The systemd wrapper is a decent start for keeping the process alive, but as others have mentioned, it doesn't address the core dependency on that API being responsive.

For network hiccups, I'd layer in a quick retry with a short timeout, but also log a local timestamp on every successful poll. Another cron job just checks that timestamp file's age. If it's stale, that triggers a separate, simple notification via a different method (like a curl to a totally separate service). It's a bit clunky, but it means you'll at least know your window is broken, even if you can't see through it.



   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

While the logic of querying the log buffer for authentication failures is sound, using the buffer as your primary source creates a retention risk that undermines the entire alerting purpose. That buffer is volatile; during a sustained attack, the very events you want to alert on can be evicted before your next poll.

You're already using the API, so shift your query to target the firewall's log forwarding streams or its queryable log database directly (using `type=log` with the appropriate log-type filter). This pulls from the same durable store that feeds your syslog aggregator, giving you consistency and eliminating the blind spot of the rolling buffer. The event IDs and structure are identical.

The thresholding and alerting logic remains valid, but now it's built on a stable data foundation.


SQL is not dead.


   
ReplyQuote
(@amyw)
Honorable Member
Joined: 2 months ago
Posts: 427
 

Glad you tackled this! I've used the same API method for PAN-OS to monitor admin user logins (successful ones) and it's surprisingly fast.

But I found the log buffer a bit flaky for this. A week of heavy traffic logging pushed my auth events out before the next poll. Switched to using `type=log` with a filter for `logtype eq 'system'` and never lost one again. Same event IDs, just more reliable.

You could keep the same threshold logic, just point it at the persistent store.


measure twice, ship once


   
ReplyQuote
(@infra_ops_guru)
Honorable Member
Joined: 6 months ago
Posts: 397
 

That's the pragmatist's version of a quorum system. You're acknowledging the tautology and building a second, simpler witness. The reliability often comes from that very simplicity - fewer moving parts to fail.

I'd push back slightly on it just moving the problem. Yes, the VPS is another dependency, but it's a *different* one. You've diversified failure domains. The attacker now needs to compromise both the firewall's management API *and* your out-of-band VPS's connectivity to blind you completely. That's a meaningful, if not perfect, improvement over a single point of failure.

The real trick is keeping that ping truly minimal. The moment you add logic to it, you start needing to monitor that logic.


infrastructure is code


   
ReplyQuote
(@bluefox)
Reputable Member
Joined: 3 months ago
Posts: 228
 

Yep, that "tall ladders" line is the perfect visual for it. I think the vendor demos miss the real point - they're selling you a window into the house, but the window frame is part of the same wall that's on fire.

Your out-of-band path is the right answer, but you're spot on about the cost. Most teams won't fund a separate bastion just for a "what if" monitor. So we end up with these turtles, pretending systemd is a foundation.



   
ReplyQuote
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
 

You're right about the cost barrier, but that makes the statistical risk assessment even more critical. The decision isn't purely technical, it's about calculating the mean time between failures of your primary monitoring path versus the cost of a second dependency.

Most teams accept the risk because the probability of total API blindness occurring before they'd notice through other channels feels low. The fallacy is assuming that failure mode is independent of an attack. A targeted actor will exploit that exact blind spot simultaneously.

So the real question becomes: what's the financial or operational impact of being blind for, say, twelve hours? If it's high, the VPS cost is trivial. The problem is we rarely quantify it that way.


prove it with data


   
ReplyQuote
Page 1 / 2