Hello everyone. I’ve been evaluating QRadar for the past few months, primarily to help our manufacturing operations meet several compliance frameworks, with an immediate priority on PCI-DSS due to our integrated B2B ecommerce platform. Coming from a background in ERP and inventory systems, I’m accustomed to building very specific, auditable reports, and I wanted to apply that same detail-oriented approach here.
I’ve just finished a rather involved process of creating a dedicated PCI-DSS compliance dashboard by building a set of custom rules and log sources. The goal was to monitor controls like failed access attempts to cardholder data environments, changes to user privileges, and firewall rule modifications. While the documentation is thorough, I found the practical steps for moving from a rule test to a stable dashboard item required quite a bit of piecing together.
I wanted to share my steps and get feedback on whether this approach is optimal, as I am cautious about potential blind spots in my logic. Below is a summary of the workflow I followed for one specific control (monitoring failed administrative login attempts). I have taken several screenshots which I can link to if that is permitted.
First, I identified the relevant log sources (Windows Event Logs from our domain controllers and firewall syslogs). I created a custom log source extension to normalize certain fields that weren't being parsed by the default DSM, specifically for our legacy application servers. This was necessary to ensure the events were categorized correctly under the "Authentication Failed" category.
Next, I built a custom rule. I set the test mode to "On" and used a relatively long test period to gather enough data. The rule criteria were focused on events where the username belonged to a privileged group (which I defined in a reference set) and the outcome was failure. I also added an exception for known service account lockouts due to scheduled tasks, which I managed through a separate reference set that is updated weekly.
After the rule was tested and deployed, I created a custom AQL query to further refine the data for the dashboard. This query joined the rule events with asset information to pinpoint the exact servers involved. Finally, I added this query as a dashboard item using a bar chart to show trends and a table listing the most frequent offending usernames and source IPs.
My primary question for the community is about maintaining these reference sets efficiently. Is there a best practice for automating the population of exception reference sets, perhaps via a scheduled script that pulls from our NetSuite employee database or Active Directory? Also, I am curious if anyone has structured their PCI dashboard to directly map dashboard items to specific PCI-DSS requirements (like 8.1.1 or 10.2.1) and how you manage updates when the rules evolve. I am concerned about the audit trail of changes to the rules themselves over time.
The transition from a tested rule to a stable dashboard item is often the most fragile part of the process. You're right to be cautious about blind spots, especially with log source reliability.
For administrative login attempts, a common oversight is not accounting for the normalization of different log formats from your various systems. Your rule might be built on a specific QID map, but if a new firewall model sends a slightly different syslog field, events can silently drop. I'd recommend building a secondary, broader correlation rule that triggers on any high-severity authentication event from your cardholder data environment network groups, then use that as a failsafe to alert you when your primary, precise rule stops firing.
Could you elaborate on how you're handling the time window and aggregation for these attempts? PCI-DSS Requirement 8.1.6 implies a need for near-real-time detection of pattern-based failures, not just daily aggregates.
Every dollar counts.
Your point about the secondary, broader rule is spot on and something I enforce in any monitoring system. I call it a "rule health check." If your primary correlation rule for, say, "10 failed admin logins in 5 minutes" fires zero times for 48 hours while your broader "any authentication failure from CHDE" rule is still active, you have a detection gap that needs immediate investigation, not a quiet success.
Regarding the time window for Requirement 8.1.6, you can't rely on daily aggregates. The detection has to be near-continuous. I implement this by setting the correlation rule to run on a sliding window, typically 5 to 10 minutes, and triggering on count thresholds within that window. The real trick is ensuring your event pipeline latency is low enough that events are normalized and available for correlation within that short window. If your log sources have a 3-minute processing delay, a 5-minute window is already broken.
What's your method for testing the real-world latency of events from source to the point they're available for a rule? I've seen teams burn hours because they assumed sub-second ingestion only to find a queue was backing up.
Show me the benchmarks