I've noticed several new members expressing confusion about Panther's data model. This is a critical concept, as it forms the foundation for all your detection logic. In simpler terms, the data model is a layer of abstraction that normalizes raw log data from different sources into a consistent schema.
Think of it this way: a login event from AWS CloudTrail and a login event from Okta have vastly different raw JSON structures. Panther's data model allows you to map both to a unified `Login` schema. This means your detection rules can be written against the normalized data model, not the raw source-specific format.
A straightforward example is mapping authentication logs. Consider these two simplified log sources:
* **Okta SystemLog Event (Raw):**
`{"eventType": "user.session.start", "actor": {"alternateId": "[email protected]"}, "client": {"ipAddress": "192.168.1.1"}}`
* **AWS CloudTrail Event (Raw):**
`{"eventName": "ConsoleLogin", "userIdentity": {"userName": "[email protected]"}, "sourceIPAddress": "192.168.1.1"}`
In a Panther data model, you would define a common structure, perhaps called `AuthenticationEvent`, with fields like:
* `p_event_type` (normalized from `eventType` or `eventName`)
* `actor_user_id` (normalized from `actor.alternateId` or `userIdentity.userName`)
* `source_ip` (normalized from `client.ipAddress` or `sourceIPAddress`)
Your detection rule then runs on the normalized `AuthenticationEvent`. You write a single rule to check for `source_ip` from a known malicious geo-location, and it will analyze both Okta and AWS logs correctly. The primary benefits are:
* **Simplified Rule Writing:** You don't need separate logic for each log source.
* **Consistency:** A field like `actor_user_id` means the same thing everywhere.
* **Maintainability:** Adding a new data source (e.g., GCP logs) only requires extending the data model mapping, not rewriting all related rules.
Without this layer, you'd be writing complex, source-specific parsing logic within every single detection, which is error-prone and difficult to scale.
prove it with data
Great example. That's the exact "aha" moment for most people. Your example shows the **why** perfectly.
One thing I'd add from practice: the real power comes when you want to add a third source, like a G Suite log. You just map it to your existing `AuthenticationEvent` model, and *all* your existing detection rules for logins instantly work on the new data source without any changes.
Saves a ton of time compared to writing rules for each source format.
—b