App-ID delay was consistent baseline overhead, not jitter. We measured +8-12ms across >95% of inspected packets for known apps like HTTPS/SQL. The jitter came from the GWLB's burst credits, as others noted.
On log volume, the tipping point wasn't daily GB but peak ingestion rate. Our Panorama VM (m5.4xlarge) hit 100% CPU on the log collector process during 30-minute windows where we ingested over 15,000 logs/sec. Daily average of 180GB was fine, but the spikes killed query performance for hours. That's what pushed us to S3.
Trust, but verify
That 15,000 logs/sec peak ingestion metric is a useful benchmark. It's a good reminder that sizing for the average is a waste of time.
We hit a similar wall, and our "solution" was the same: ditch Panorama for historical logs and ship everything to our existing SIEM. The operational overhead is real, but at least the queries run.
Your GWLB credit issue is the real architectural flaw. Once you exempt those high-volume flows, you're just paying for a very expensive packet forwarder.
Sizing for peak ingestion rate is the only way to avoid that performance wall. The 15,000 logs/sec threshold you both mention aligns with what I've seen in our own logging load tests against the Panorama VM; the collector process is the clear bottleneck.
Your point about the expensive packet forwarder is accurate. Once you start carving out exemptions for performance, the cost-per-inspected-packet metric goes way up. We actually measured this: the effective cost for inspecting the remaining "risky" east-west traffic was nearly 3x the initial projected cost based on total flow volume.
Did your SIEM integration handle the native Palo Alto log format directly, or did you need to normalize the data first? That became a secondary overhead for us.
BenchMark