Skip to content
Notifications
Clear all

Just built a simple dashboard to track Cybereason alert trends in Grafana

7 Posts
7 Users
0 Reactions
29 Views
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
Topic starter   [#14555]

Having recently completed a significant deployment of Cybereason across our enterprise endpoints, we were immediately confronted with the classic operational challenge: the platform generates a high-fidelity stream of security alerts, but deriving actionable intelligence and identifying macro-trends from the native management console proved cumbersome for our SOC analysts and leadership. The need was for a historical, at-a-glance view of alert volumes, severities, and categories over time, integrated into our existing operational dashboards.

To address this, I've constructed a purpose-built Grafana dashboard that visualizes Cybereason alert data by querying its REST API. The architectural flow is a straightforward ETL pipeline: a Python-based collector fetches data at regular intervals, transforms it into metrics, and pushes them into Prometheus, which Grafana then queries. The key was structuring the API response into meaningful time-series dimensions.

The collector script, scheduled via cron, performs authenticated calls to the `/rest/visualsearch/query/v1` endpoint with a carefully crafted query to aggregate alerts from the past hour. Here is the core of the data extraction logic:

```python
import requests
import json
import time
from prometheus_client import Gauge

# Prometheus Gauges
alert_total = Gauge('cybereason_alerts_total', 'Total alert count', ['severity', 'category'])
alert_by_machine = Gauge('cybereason_alerts_by_machine', 'Alerts by machine', ['machine_name', 'severity'])

def fetch_alert_metrics():
api_url = "https://.cybereason.net/rest/visualsearch/query/v1"
headers = {"Content-Type": "application/json"}
query = {
"queryPath": [{"requestedType": "Alert", "filters": [], "isResult": True}],
"totalResultLimit": 1000,
"perGroupLimit": 1000,
"perFeatureLimit": 100,
"templateContext": "ALERTS",
"customFields": ["severity", "category", "machineName"]
}
response = requests.post(api_url, headers=headers, json=query, auth=('', ''))
data = response.json()

# Reset gauges to handle disappearing alerts
for metric in [alert_total, alert_by_machine]:
for metric in metric._metrics.values():
metric.clear()

for alert in data.get('data', {}).get('results', []):
sev = alert.get('severity', 'UNKNOWN')
cat = alert.get('category', 'UNKNOWN')
machine = alert.get('machineName', 'UNKNOWN')

alert_total.labels(severity=sev, category=cat).inc()
alert_by_machine.labels(machine_name=machine, severity=sev).inc()
```

The resulting dashboard panels include:
- A time-series graph showing alert count per severity (Critical, High, Medium, Low) over the last 7 days.
- A stacked bar chart breaking down alerts by category (Malware, Suspicious Activity, Exploit, etc.) per day.
- A table showing current alert hotspots by endpoint hostname, sorted by Critical alert count.
- A stat panel showing the percentage change in total alerts compared to the previous week.

Initial results have been illuminating. We've identified a previously unnoticed correlation between a specific software update (tracked in a separate system) and a spike in "Suspicious Activity" alerts every Tuesday morning, which has allowed us to create an exclusion rule and reduce noise by approximately 18%. The integration also provides a shared view for both security and IT operations teams, bridging a visibility gap.

Potential pitfalls to consider:
- The Cybereason API can throttle under heavy load, so aggressive polling intervals are not advised.
- The metric cardinality can explode if you label by every possible dimension; I recommend limiting labels to core attributes like severity, category, and machine.
- Authentication credentials must be securely managed, preferably via a vault, as the script requires significant privileges.

This approach transforms the alert data from a reactive list into a quantifiable, trend-able metric. I am considering expanding the pipeline to feed raw alert logs into a data lake for longer-term retention and more complex correlation analysis outside the Cybereason platform's native retention period.


Data is the source of truth.


   
Quote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

This is a fantastic approach. So many security tools have great data but poor historical visualization, forcing you to build exactly what you need externally. Using Prometheus as the metrics layer is a smart, scalable choice.

One thing I'd be curious about is how you're handling the query logic for `visualsearch/query/v1`. Crafting those queries can be finicky. Did you find the biggest challenge was filtering out the noise to get clean aggregates, or something else like pagination or rate limits?


Keep it civil, keep it real


   
ReplyQuote
(@devops_shift_worker)
Reputable Member
Joined: 4 months ago
Posts: 290
 

Nice. We did something similar a while back, but hit a snag you might run into later. The `visualsearch/query/v1` endpoint can get really chatty if you're pulling all alerts over a long period - watch out for those API rate limits during peak breach-fest hours.

We ended up adding a Redis cache for the raw query results (just for 5 minutes) before the transform/prometheus push. Took the load off Cybereason when our cron intervals overlapped. Also, the query JSON itself becomes a nightmare to maintain. We templated it with Jinja2.

What are you using for alert grouping in Grafana? We found the default Prometheus histograms were okay, but cardinality exploded on the `category` label. Had to add a filter to consolidate the "low and clear" stuff into an `other` bucket.


NightOps


   
ReplyQuote
(@db_diver)
Reputable Member
Joined: 7 months ago
Posts: 333
 

The approach of structuring the raw API response into proper time-series dimensions before the Prometheus push is the critical piece so many people miss. They'll dump the raw JSON as a gauge and try to wrangle it in Grafana, which becomes unmaintainable.

My one addition would be to consider your data retention and aggregation strategy from the start. Prometheus is great for the recent time window, but if you're looking at trends over weeks or months, you'll likely need downsampling. I'd recommend having your collector also write to a separate, cost-effective time-series store like ClickHouse or even a managed offering like TimescaleDB for those longer-term, aggregated views. This keeps your Prometheus instance lean for operational dashboards while still supporting historical analysis.

What's your plan for storing and querying alert data beyond the standard Prometheus retention period?


SQL is not dead.


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Agreed on structuring before the push. That cardinality explosion is a real cost driver in Prometheus.

We skip long-term storage in Prometheus entirely. The collector writes two streams: summarized metrics for Prometheus (last 15 days), and raw, de-normalized alert documents to S3 as JSONL. For long-term trends, we use Athena to query the S3 bucket. It's cheaper and we can do ad-hoc aggregations they'd never think to ask for initially.


Trust, but verify


   
ReplyQuote
(@jenniferg)
Estimable Member
Joined: 3 months ago
Posts: 76
 

Structuring the API response into dimensions before the push is absolutely the right move. It forces you to think critically about your data model from the start. My only addition to that point is to also formalize a schema for those dimensions early on, maybe even version it. As your dashboard evolves and you add new visualizations, that consistent structure prevents drift and makes it much easier for a colleague to understand or take over the project later.


Let's keep it real.


   
ReplyQuote
(@benchmark_bob_43)
Reputable Member
Joined: 5 months ago
Posts: 243
 

Pushing to Prometheus is the right call, but I hope you're not just pushing a raw count. The real power is in the dimensions you expose - are you capturing things like `detection_type`, `machine_os`, or `malop_status` as labels? That's where you start spotting the real trends, like a spike in "malicious PowerShell" alerts from engineering workstations.

And for the love of benchmarks, please tell me you're not running that cron every hour with a static time window. The collector should use the last successful run's timestamp as the `startTime` in the query, otherwise you're going to have gaps or duplicates during any collector hiccup. I learned that one the hard way. 😅



   
ReplyQuote