Skip to content
Notifications
Clear all

Showcase: My custom alert for suspicious Azure service principal activity.

3 Posts
3 Users
0 Reactions
31 Views
(@integration_ian_3)
Honorable Member
Joined: 4 months ago
Posts: 411
Topic starter   [#17869]

Hey everyone! I've been living in the Cortex XDR API and webhook configuration for the better part of six months, and I wanted to share a custom alert I built that's been incredibly useful for our cloud environment. We're heavy Azure users, and while XDR's native cloud alerts are good, I found I needed something more specific to catch suspicious Service Principal activity that might fly under the radar.

The goal here was to detect actions like new credential additions, high-risk consent grants, or authentication from unexpected locations for a service principal—classic precursors to a supply chain attack or lateral movement. I'm using Make (formerly Integromat) as my middleware to keep things flexible, but this could be done with Zapier, a Python script, or even XDR's own native webhook actions.

Here's the basic workflow architecture:
1. **Azure Monitor (Diagnostic Settings)** streams Azure AD audit logs to an **Azure Event Hub**.
2. A small **Azure Function** (Node.js) filters for only `ServicePrincipal` events and transforms the payload.
3. The Function posts the filtered high-risk events to a **Cortex XDR Webhook** (from the `Create an incident` action).
4. XDR creates an incident, which I then enrich with some internal API data for context.

The real magic is in the Azure Function filtering logic. I'm not just grabbing all events; I'm looking for very specific `OperationName` values.

```javascript
// Example of the key filtering logic in the Azure Function
const highRiskOperations = [
'Add service principal credentials', // Credential stuffing
'Add delegated permission grant', // Risky OAuth consent
'Add app role assignment to service principal', // Permission elevation
'Sign-in activity' // But only from unexpected geolocations
];

if (highRiskOperations.includes(eventBody.OperationName)) {
// Enrich the event with additional context from Azure Graph API
enrichedEvent = await enrichWithServicePrincipalDetails(eventBody);
// Forward to Cortex XDR webhook
await postToXDR(enrichedEvent);
}
```

**Some key gotchas and lessons learned:**

* **Rate Limiting:** The XDR webhook endpoint can throttle you. I implemented a simple exponential backoff in my Function after hitting 429s. It's crucial for burst scenarios.
* **Payload Mapping:** The XDR incident field mapping (like `title`, `severity`, `custom_fields`) needs careful attention. I map the Azure operation to a severity score—"Add credentials" is a `high`, while a suspicious sign-in might be `medium`.
* **Alert Fatigue:** Tune the filters aggressively. My first version created too much noise. I added a whitelist for known, secure service principals used by our CI/CD pipelines.
* **Enrichment is Key:** The raw Azure event gives you a service principal ID. I added a second API call (from within XDR using a playbook) to fetch the app's display name and owner details. This makes the alert immediately actionable for our SOC.

This integration has let us catch a couple of misconfigured automation scripts and one genuinely suspicious credential dump attempt. The beauty of using XDR as the consolidation point is that these alerts now correlate with other endpoint and network data we feed into it.

I'm curious if anyone else has built similar cloud-focused traps? How are you handling the authentication for the various APIs (Managed Identity for Azure, API keys for XDR)? I've got some thoughts on securing those handoffs if anyone's interested.

-- Ian


Integration Ian


   
Quote
(@chrism)
Reputable Member
Joined: 3 months ago
Posts: 326
 

Love this approach! Using a serverless function as the filter before hitting the webhook is smart - keeps your alerting pipeline lean and cost-effective.

I've done something similar but with a Log Analytics query as my trigger, feeding into a Logic App. Found I had to be really careful about the "unexpected locations" logic to avoid noise from legitimate CI/CD pipelines spinning up in different regions.

Any specific thresholds you're using for what counts as "high-risk" consent grants? That's an area where we've had to tune ours pretty aggressively.


K8s enthusiast


   
ReplyQuote
(@integration_ian)
Honorable Member
Joined: 5 months ago
Posts: 396
 

Noise from CI/CD regions was a big problem for me too. I added a three-step filter that's mostly stopped the false positives:
- Allow list of known cloud provider regions our pipelines actually use.
- Exclusion for service principals with a specific naming convention (prefix "CICD-").
- A time window check. If the first auth from a new region happens between 2-5 AM local time to the SP's usual base, it gets a critical severity.

>Any specific thresholds for "high-risk" consent grants

We treat any consent grant to an app we haven't seen before in our tenant as high-risk, full stop. That triggers a manual review. The bigger tuning issue was dealing with admin consent vs user consent. We eventually flagged any user consent that required more permissions than they already had for that app.


Integration is not a project, it's a lifestyle.


   
ReplyQuote