We had a gap. Cato's SASE platform threw security alerts, but they weren't hitting our on-call rotation in PagerDuty. Had to close that loop.
Here's the webhook integration we built. Cato's Event Forwarding sends to a small AWS Lambda, which formats and pushes to PagerDuty.
Key Lambda function (Python):
```python
import json
import os
import requests
PD_EVENTS_URL = "https://events.pagerduty.com/v2/enqueue"
INTEGRATION_KEY = os.environ['PD_INTEGRATION_KEY']
def lambda_handler(event, context):
for cato_alert in event.get('alerts', []):
payload = {
"routing_key": INTEGRATION_KEY,
"event_action": "trigger",
"payload": {
"summary": f"Cato: {cato_alert.get('eventType', 'Alert')}",
"severity": map_severity(cato_alert.get('severity')),
"source": cato_alert.get('sourceName', 'Cato'),
"custom_details": cato_alert
}
}
requests.post(PD_EVENTS_URL, json=payload)
return {"statusCode": 200}
def map_severity(cato_sev):
sev_map = {"Critical": "critical", "High": "error", "Medium": "warning", "Low": "info"}
return sev_map.get(cato_sev, "info")
```
Cato side: In the management console, set up Event Forwarding to this Lambda's API Gateway endpoint. Filter for the alert types you need (e.g., Threat Prevention, Anomaly).
Lessons:
* The Lambda gives you a place to filter noise before PagerDuty.
* Map Cato's severities to PD's four levels explicitly.
* Test with real low-severity alerts first.
cg
YAML all the things.
Good approach with the webhook proxy. Did you consider any retry logic in the Lambda? The PagerDuty events API can have intermittent failures.
Our team also added a deduplication step using the Cato alert ID as the `dedup_key`. It prevents duplicate incidents for the same event. PagerDuty handles that nicely if you include it in the payload.
Great point about retry logic - it's easy to miss that. The AWS Lambda runtime will retry on its own for asynchronous invocations (like from an SQS queue), but if you're triggering directly from something like EventBridge, a failure just drops the event.
We added a simple retry decorator for the PagerDuty API call, but honestly, the `dedup_key` tip you mentioned was even more valuable for us 😄. We had a bug where our forwarding rule in Cato fired twice on the same event. Without that key, we'd have created duplicate incidents and confused the on-call engineer.
Our mapping looks something like this now in the payload:
```python
"dedup_key": cato_alert.get('id'),
"payload": {
"custom_details": cato_alert
}
```
Lets PagerDuty handle the deduplication perfectly.
Clean code is not an option, it's a sanity measure.
Thanks for sharing the Lambda code. That's really helpful for seeing the full flow. One thing I noticed: the `requests.post` call doesn't have any error handling. What happens if the API call fails? Does the Lambda just return 200 anyway?
Also, I'm curious about the `map_severity` function. Have you run into any Cato alerts with a severity you didn't map, and they defaulted to "info"?
Great example of the mapping function, and yeah, the default to "info" is exactly right. We used that same approach, but hit a snag where some of our internal test events from Cato didn't have a severity field at all - it was `null`. Our map function choked on that and the whole Lambda failed. Had to add a guard clause like `if not cato_sev: return "info"` before the dictionary lookup.
On the error handling point, that's a real issue. The Lambda returning 200 after a failed `requests.post` means you'd lose the alert. At a minimum, you should check the response status and raise an exception to let Lambda's built-in retry kick in, if you've configured it. Even better, wrap the call in a try-except and send to a dead-letter queue for manual review on repeated failures.
You're absolutely right about the value of the deduplication key. We initially overlooked it too, and found it was even more critical than retry logic in our specific case because of an unexpected behavior in Cato's event forwarding. Certain rule configurations could, under high load, apparently generate two identical webhook payloads milliseconds apart. Without the dedup_key, that created noise. The retry logic handles transport failures, but the dedup_key handles what I'd call "source duplication," which is a different class of problem.
Implementing the dedup_key actually simplified our error handling approach. Since PagerDuty will accept and then deduplicate a second identical event, we became less concerned about the edge case of a retry accidentally creating a duplicate incident. It allowed us to focus our retry logic purely on network or API availability issues, not idempotency.
Let's keep it constructive
Good point about source duplication being a different class of problem. But you're placing a lot of faith in the dedup_key solving all your noise issues. What about alerts that are slightly different? A second webhook for the same event might have a timestamp that varies by a millisecond, and your dedup logic might not catch it if you're just using the base alert ID.
Also, focusing your retry logic purely on network issues because of the dedup key is a risky trade-off. You're assuming the second, deduplicated event will always make it through. If your network is flaky, you could lose the original and its duplicate, leaving you blind. The dedup key helps with clean data flow, but it doesn't replace the need for a solid failure-handling policy that includes queueing or storing state.
Trust but verify.
Oh, that's a good distinction about Lambda retries. I hadn't thought about the difference between async (like SQS) and direct triggers. So if it's coming from EventBridge and fails, the event is just gone?
The dedup_key using the alert ID is super clever. Does PagerDuty ever reject a dedup_key if it's too long or has special characters? I'd be worried about that.