Skip to content
Notifications
Clear all

Where do I start with logging? The volume is overwhelming.

1 Posts
1 Users
0 Reactions
32 Views
(@elliotn)
Reputable Member
Joined: 3 months ago
Posts: 291
Topic starter   [#14068]

I’ve deployed a Palo Alto Networks NGFW series (specifically a PA-3260) in a hybrid cloud environment, and while the security capabilities are robust, the sheer volume of generated logs is creating a significant data engineering challenge. My initial setup forwards all traffic, threat, and system logs to a central SIEM, resulting in approximately 1.2 TB of raw log data daily. Without a deliberate filtering and enrichment strategy, this volume is not only costly to store but makes meaningful analysis nearly impossible.

Based on my benchmarks, a methodical, layered approach to log management is critical. I propose starting with a clear data pipeline strategy before enabling any log forwarding. The primary goal is to reduce noise and increase signal by several orders of magnitude.

**Phase 1: Source-Side Filtering & Sampling (On the NGFW Itself)**
This is the most efficient place to reduce volume. Do not log everything.
* **Traffic Logs:** Apply aggressive filters. Exclude known-low-risk traffic first. For example, create a rule to omit logging for internal health checks, trusted SaaS provider IP ranges, or specific low-security zones.
* **Threat Logs:** Focus on severity. Consider disabling logging for "informational" severity threats initially. Ensure logging for "medium" and above is on, but review the profiles to avoid duplicate alerts.
* **Implement Sampling:** For very high-volume, non-critical traffic (e.g., general web browsing from a specific subnet), you can use session sampling. A 1:1000 sample can still provide traffic flow visibility without the burden.

**Phase 2: Structured Parsing & Enrichment (In the Pipeline)**
Raw syslog is inefficient. Parse and structure logs immediately upon ingestion.
```python
# Example conceptual pipeline step using a tool like fluentd or logstash
filter {
if [type] == "pan-firewall" {
grok {
pattern => "%{SYSLOG5424PRI}%{NONNEGINT:log_version} %{TIMESTAMP_ISO8601:generation_time} %{HOSTNAME:device_name} %{WORD:log_type} %{DATA:log_subtype} %{NONNEGINT:serial} %{DATA:source_ip} %{DATA:destination_ip} %{GREEDYDATA:rest}"
}
# Add enrichment: geoip for src/dst, threat intel lookup, internal asset tagging
translate {
field => "[destination_ip]"
destination => "[internal_asset_owner]"
dictionary => { "10.0.1.5" => "database-server-prod", "192.168.10.10" => "file-server-01" }
fallback => "unknown"
}
}
}
```
This transformation turns a text blob into indexed fields, enabling efficient querying and reducing storage by discarding unnecessary raw text after parsing.

**Phase 3: Tiered Storage & Retention Policy**
Not all logs need hot storage. Define clear lifecycle rules.
* **Hot (7 days):** All threat logs (medium+), traffic logs for critical assets/demilitarized zones, all configuration logs.
* **Warm (30 days):** Sampled traffic logs, all URL filtering logs.
* **Cold (1 year):** Aggregated flow data (source/dest/port/bytes), compressed system logs for compliance.

**Key Metrics to Establish Baseline:**
Before applying filters, measure for one business day:
1. Logs per second (LPS) by log type (traffic/threat/system).
2. Average log message size in bytes.
3. Top 10 source/destination IP pairs by log volume.
4. Percentage of logs with severity "informational" or "low."

These metrics will identify the largest contributors to volume and allow you to measure the impact of each filtering rule. Start by implementing Phase 1 rules incrementally, monitoring the reduction in LPS for each change. The objective is not to eliminate logs, but to architect a manageable stream where critical security events are not drowned in operational noise.

-- elliot


Data first, decisions later.


   
Quote