Right about the buffer volatility. The `type=log` source is the move.
But pulling everything into your main poll can bloat that operation if you're collecting a lot of traffic data. You're trading the buffer risk for a potential performance hit on your primary collection job. A separate, dedicated query to the persistent log source with a tight filter is often cleaner.
That's a good point about bloating the primary poll. Would a separate query with a `log-type eq 'system'` filter still hit performance if you're just doing something like `last-pulled-at > timestamp`, or does the firewall have to process the whole log set to find matches anyway?
Nice work getting ahead of that visibility gap. The log buffer approach works for a proof of concept, but as others have hinted, you're gonna want to switch to the persistent log query soon. I had a nearly identical script running for about a month before a busy Friday afternoon filled the buffer and pushed out a bunch of failed attempts before the next poll. I missed a whole series of probes.
The `type=log` query is your friend here. Just a heads up though, depending on your log volume, you might want to schedule that separate from your main traffic log collection. I found running a filtered query just for system logs around the clock worked without a hitch, but yanking everything into one job caused some timeouts.
it worked on my machine
Preach. But quantifying that impact is a political and budgeting exercise, not a technical one.
You'll get laughed out of the room for asking for $20/mo for a VPS "in case the management plane is under attack" when the primary monitoring costs $50k a year. The cost isn't trivial to the team signing the PO, it's another line item requiring justification. So they accept the hidden, unquantified risk instead.
Your stack is too complicated.
You're right about the absurdity, but the vendor demos are even worse than just being a tautology. They're selling you a system that alerts *itself* about its own failure, when the mechanism to transmit that alert is part of the same failed component.
The real-world analogy that grates on me is a burglar alarm that requires the home's main power to call the police. If they cut the line, the alarm can't even report the outage. Your out-of-band path is like a cellular backup, but as you said, that's a separate project with its own BOM and maintenance. So we all just hope the power line is buried deep enough.
Show me the benchmarks.
Good on you for tackling the gap directly. That local buffer polling will bite you, though. It's volatile and sized for a couple minutes of logs, not security monitoring. As user238 said, you'll miss probes the first time the box sees a traffic spike.
Switch to a `type=log` query with a `log-type eq 'system'` filter as soon as you can. It hits the same disk-backed logs that syslog would. Just make it a separate, dedicated job with a tight time filter - don't tack it onto whatever you're using to pull traffic logs, or you'll get contention and timeouts.
And for the love of all that's holy, don't just alert on threshold. Log the raw failed attempts with timestamps somewhere durable outside the firewall. If you're getting brute-forced, you'll want the full timeline for the incident report, not just a count that tripped your trigger.
Good call on the API key in an env var. I've seen way too many of these scripts commit that straight to git.
Your buffer query will work, until it doesn't. The `subtype eq authentication` filter you're using is correct, but as others said, you need to switch to `type=log` and query the system logs specifically. The buffer will rotate out from under you during a traffic surge, which is ironically when you'd most want this alert.
One more thing on the threshold: consider adding a separate, lower threshold for distinct usernames from the same IP, not just total attempts. That can catch smarter, slower credential stuffing that might fly under a simple attempt count.
YMMV
That's a smart way to fill the gap quickly. I'm trying to build something similar for our setup, but I'm new to the PAN-OS API. Could you share a bit more about how you handled the `last-pulled-at > timestamp` logic for the rolling window? I'm worried about missing entries if the script restarts.