Skip to content
Notifications
Clear all

Newbie struggling with the 'data model' concept. Any simple examples?

2 Posts
2 Users
0 Reactions
42 Views
(@bookworm)
Reputable Member
Joined: 3 months ago
Posts: 281
Topic starter   [#15544]

I've noticed several new members expressing confusion about Panther's data model. This is a critical concept, as it forms the foundation for all your detection logic. In simpler terms, the data model is a layer of abstraction that normalizes raw log data from different sources into a consistent schema.

Think of it this way: a login event from AWS CloudTrail and a login event from Okta have vastly different raw JSON structures. Panther's data model allows you to map both to a unified `Login` schema. This means your detection rules can be written against the normalized data model, not the raw source-specific format.

A straightforward example is mapping authentication logs. Consider these two simplified log sources:

* **Okta SystemLog Event (Raw):**
`{"eventType": "user.session.start", "actor": {"alternateId": "[email protected]"}, "client": {"ipAddress": "192.168.1.1"}}`

* **AWS CloudTrail Event (Raw):**
`{"eventName": "ConsoleLogin", "userIdentity": {"userName": "[email protected]"}, "sourceIPAddress": "192.168.1.1"}`

In a Panther data model, you would define a common structure, perhaps called `AuthenticationEvent`, with fields like:
* `p_event_type` (normalized from `eventType` or `eventName`)
* `actor_user_id` (normalized from `actor.alternateId` or `userIdentity.userName`)
* `source_ip` (normalized from `client.ipAddress` or `sourceIPAddress`)

Your detection rule then runs on the normalized `AuthenticationEvent`. You write a single rule to check for `source_ip` from a known malicious geo-location, and it will analyze both Okta and AWS logs correctly. The primary benefits are:

* **Simplified Rule Writing:** You don't need separate logic for each log source.
* **Consistency:** A field like `actor_user_id` means the same thing everywhere.
* **Maintainability:** Adding a new data source (e.g., GCP logs) only requires extending the data model mapping, not rewriting all related rules.

Without this layer, you'd be writing complex, source-specific parsing logic within every single detection, which is error-prone and difficult to scale.


prove it with data


   
Quote
(@brandonj)
Reputable Member
Joined: 3 months ago
Posts: 253
 

Great example. That's the exact "aha" moment for most people. Your example shows the **why** perfectly.

One thing I'd add from practice: the real power comes when you want to add a third source, like a G Suite log. You just map it to your existing `AuthenticationEvent` model, and *all* your existing detection rules for logins instantly work on the new data source without any changes.

Saves a ton of time compared to writing rules for each source format.


—b


   
ReplyQuote