We've been running Google Chronicle's forwarder VMs (the on-premise ingestion component) in our production environment for approximately nine months. The deployment is on Google Compute Engine, using the recommended `n2-standard-4` machine type. Over the last several weeks, we've observed sustained CPU utilization averaging between 75-90%, with frequent spikes to 100%. This is occurring even though our log volume is consistent and well within the documented throughput limits for this instance size.
Our initial troubleshooting has ruled out obvious external factors. The issue persists across multiple forwarders in a load-balanced group. We are not ingesting unusually large or malformed files, and the source systems (a mix of on-premise syslog and cloud Pub/Sub topics) show no signs of a traffic surge. The high CPU load appears to be intrinsic to the forwarder processes themselves.
We've conducted a standard diagnostic profile and examined the forwarder configuration. The relevant section of our `chronicle-forwarder.conf` is below:
```yaml
buffer:
disk_buffer:
path: /var/spool/chronicle-forwarder
max_buffer_size_mb: 102400
memory_buffer:
max_buffer_size_mb: 512
processing:
batch_size: 1000
batch_timeout_millis: 1000
compression: SNAPPY
sources:
- syslog:
port: 10514
protocol: TCP
max_connections: 500
```
Our primary observations and questions for the community are:
* Is this sustained high CPU a known characteristic of the forwarder, particularly with the Snappy compression enabled? We are considering testing without compression, though this would increase network egress costs.
* Could the `batch_timeout_millis` setting be too aggressive, causing excessive processing cycles even when the batch size isn't full?
* Has anyone performed meaningful benchmarking on forwarder VM sizing? The official sizing guide seems to focus on throughput capacity (MB/s) but is less clear on steady-state CPU resource consumption. We are contemplating a move to `n2-standard-8`, but this feels like addressing a symptom rather than the root cause.
* Are there specific `journald` or OS-level tuning parameters (we are using the provided COS image) that have proven effective in reducing system overhead? We've already adjusted the `vm.swappiness` parameter.
We are currently correlating the CPU usage with internal forwarder metrics via the monitoring API, but anecdotal evidence from others running similar volumes would be invaluable. Any insights into configuration optimizations or known performance bugs in specific forwarder versions would be greatly appreciated.
Data over dogma