Hey folks! 👋 I've been working a lot lately on bridging the gap between our GRC team's risk management needs and the actual engineering teams who own the services. We all know the classic problem: a risk dashboard gets built in ServiceNow GRC, but it's in a different system with different logins, so devs never look at it.
So I set out to create a **single-pane-of-glass dashboard in Datadog** that pulls in ServiceNow GRC data (via API) and mixes it with live operational metrics. The goal? Make tech risk visible in the same place devs are already checking for performance issues. Here's a step-by-step of my approach.
**Step 1: The Data Connection**
First, you need to get ServiceNow GRC data into your observability platform. I used Datadog's HTTP client to pull from the ServiceNow API. You'll need to create a read-only API user in ServiceNow and get the instance URL. Here's a basic Python script I run as a scheduled Lambda function to fetch "High/Critical" risks and push them as custom events:
```python
# Simplified example - fetches active tech risks
import requests
import os
instance = os.environ['SN_INSTANCE']
user = os.environ['SN_USER']
password = os.environ['SN_PASSWORD']
headers = {"Accept":"application/json"}
params = {
'sysparm_query': 'state=active^risk_score>=400',
'sysparm_fields': 'number,short_description,risk_score,owned_by'
}
response = requests.get(f"{instance}/api/now/table/sn_grc_risk",
auth=(user, password),
headers=headers,
params=params)
# ... (format and send to Datadog's API as events)
```
**Step 2: Dashboard Layout**
I built a dashboard with three key sections:
* **Top:** Summary widgets showing "Active High Risks Count" and "Avg. Time to Remediation" (pulled from GRC).
* **Middle:** A live list of active risks, including the service/team owner from CMDB mapping.
* **Bottom:** **The crucial part** – side-by-side graphs. On the left, a graph of risk score trends from GRC. On the right, key SLO/SLA metrics (error rates, latency) for the specific services flagged as high-risk. This creates an immediate "oh, this risky service is also degrading" visual connection.
**Step 3: Making it Actionable**
To get teams to actually care, I added:
* **Annotations:** When a new high-risk item is logged in GRC, it appears as a marker on the service's performance graphs.
* **Alerts:** Combined alerts that trigger not just on, say, error rate spikes, but also when a service with a high underlying risk score has any performance dip. This adds context to pagers.
* **Ownership:** Used service tags to route dashboard views and alerts directly to the responsible team's channels.
The result? Our compliance folks get their reports from ServiceNow, and our engineers see the risk context in their operational workflow. It's not just a "check-the-box" dashboard anymore.
Has anyone else tried similar integrations? I'm curious about how you're handling the mapping between GRC assets and your service catalog.
Dashboards or it didn't happen.
Pulling GRC data via a scheduled Lambda is a good start, but you're going to hit scaling and data freshness issues fast. That script fetches a point-in-time snapshot, but risks can change quickly. You'll be looking at stale data between runs.
You need to treat this as a proper data pipeline, not a cron job. Instead of a scheduled pull, set up a stream. ServiceNow can push events to an SQS queue via outbound REST messages on risk updates. Then have a Lambda consume from the queue and forward to Datadog. This gets you near-real-time updates and is far more reliable. The batch approach falls apart when you need to track risk state transitions or correlate risk creation with a simultaneous PagerDuty alert.
Also, you're only sending events. To be useful for dashboards, you need metrics. You should be emitting a gauge for `tech_risk.severity.count` tagged by service, team, and environment. That way you can graph risk trends over time and set monitors. An event log is good for a timeline widget, but you can't do SLAs or burn-down charts with it.
—davidr
Your streaming pipeline's great until you get the bill. Real-time pushes from ServiceNow, SQS, and Lambda for every risk update? That's maybe a few thousand events a day. At that volume, the cost difference is negligible. But you're right on the metrics.
Where this gets expensive is when you try to graph that `tech_risk.severity.count` gauge. If you emit it per service per environment, you're creating a new metric series for every combination. Tag explosion. Datadog's pricing is per metric per hour. A simple dashboard could cost you more than the Lambda.
Better to keep it as events and use Datadog's log analytics to generate the timeseries you need on-demand. Cheaper, and you're not paying to store redundant data.
show the math