Skip to content
Check out this Graf...
 
Notifications
Clear all

Check out this Grafana dashboard I made to track our top 10 SIEM alert sources.

2 Posts
2 Users
0 Reactions
20 Views
(@isabelm)
Estimable Member
Joined: 3 months ago
Posts: 68
Topic starter   [#13276]

As part of our quarterly compliance audit for SOC 2, I was tasked with providing a detailed analysis of our SIEM's alert provenance. While our vendor's native reporting provides aggregate volume, it lacked the persistent, at-a-glance visibility into our most volatile alert sources needed for effective configuration drift management and cost attribution. To address this, I constructed a dedicated Grafana dashboard that tracks the top 10 alert sources by volume over the last 30 days.

The dashboard is built atop our SIEM's summary index and utilizes a series of panel transformations to ensure data remains current and actionable. The core methodology is as follows:

* **Primary Data Source:** A time-series query grouping by `source_name` field, counting distinct `alert_id` over a rolling 30-day window.
* **Ranking Logic:** A combination of Grafana's built-in "Stat" panel and "Bar Gauge" panel, sorted descending, limited to the top 10 results. This is recalculated every 15 minutes via the dashboard's refresh interval.
* **Ancillary Panels:**
* A line graph showing the daily percentage contribution of the current #1 source versus all other sources, highlighting dominance shifts.
* A table listing each of the top 10 sources with exact counts and a week-over-week percentage change calculation, which is crucial for identifying emerging noise or potential misconfigurations.
* A log panel displaying the most recent 20 raw alert examples from the currently selected source, allowing for immediate pattern validation.

This visualization has already yielded several actionable insights. For instance, we identified a single network sensor responsible for 22% of all alerts due to a overly broad signature update that was not documented in the change log. Furthermore, correlating spikes in the dashboard with our release calendar allowed us to attribute a 15% increase in alerts from a specific application server to a recent deployment, a fact that was omitted from the deployment's release notes.

I am interested in discussing the metrics others track for alert source hygiene. Specifically:

* What key performance indicators, beyond raw volume, do you find most predictive of configuration drift?
* How do you structure your data pipelines to efficiently feed such dashboards without incurring excessive ingestion costs from the SIEM itself?
* For those using SOAR playbooks for alert triage, have you integrated similar source-level analytics to dynamically adjust playbook thresholds or assignment routing?



   
Quote
(@devops_rookie_2025)
Prominent Member
Joined: 4 months ago
Posts: 467
 

Hi! I'm a junior DevOps engineer at a mid-sized SaaS company (around 150 people). We run a mix of containerized services on EKS and use Loki for logs, so I see a lot of dashboards.

Since you're building on top of your SIEM's data, a few things I had to figure out the hard way:

**Refresh Rate vs. Query Cost**: Running that top-10 ranking every 15 minutes on a 30-day window could get expensive if your SIEM query charges per GB scanned. I'd check if you can aggregate the data first. We had to set up a daily summary table in Postgres for a similar dashboard to avoid surprise bills.
**Drill-Down Ability**: The native Grafana bar gauge is great for the overview. Make sure you've set up a variable or a panel link so clicking a source name can jump to a detailed log explorer view. It's a lifesaver during an audit question.
**Transformation Complexity**: If you're using a lot of "Merge" and "Filter data by value" transformations for the ranking, they can slow the dashboard load for other users. We capped similar ranking panels at 12 items and saw load times drop from ~8s to under 3s.
**Alerting on Shifts**: You might want a simple alert rule on that "daily percentage contribution" panel. If your #1 source suddenly spikes to >50% of total alerts, it could mean a misconfiguration. Setting a threshold alert was one of my first tasks, and it caught a broken log-forwarder.

For a persistent, internal tracking dashboard like this, your approach looks solid. If the goal is strictly SOC 2 evidence and not real-time response, would you consider a scheduled PDF export to a compliance folder? That way you keep the live view but also have a static artifact for the auditors.



   
ReplyQuote