Skip to content
Notifications
Clear all

Help: High CPU on PA-3220, but traffic is low. What to check?

7 Posts
7 Users
0 Reactions
39 Views
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
Topic starter   [#13568]

We've been running a pair of PA-3220s in active/passive HA for several years, and recently the active unit has been exhibiting sustained high CPU (75-90%) while the throughput is well below the platform's capacity—approximately 200 Mbps with 15,000 sessions. The passive unit sits at 5-10% CPU. This discrepancy is causing concern for stability and failover readiness.

I've ruled out the obvious: there's no ongoing threat prevention signature update, no logging flood to an external service, and the dataplane traffic figures from the CLI align with the GUI. The high load is primarily coming from the management plane (`mp` process) and the dataplane (`dp` process) in roughly equal measure, not from a single process like `wsf` or `flowd`.

So far, I have checked the following:
* Confirmed HA sync is healthy and there are no config mismatches.
* Verified session count is within specification and session setup rate is normal.
* Disabled all logging to Panorama and external syslog servers temporarily, with no impact on CPU.
* Reviewed the output of `show running resource-monitor` and `show running resource-monitor minute`—no single zone, rule, or user is generating disproportionate sessions.

The current software version is PAN-OS 10.1.6. I am considering the next diagnostic steps, but I want to be systematic. My specific questions are:

1. Beyond the standard resource monitor, what are the most useful CLI commands to pinpoint management plane CPU consumption? For example, should I be deeply inspecting `debug` commands or specific `show` commands for the internal message bus or job queues?
2. Given that the load is shared between `mp` and `dp`, could this indicate an issue with hardware acceleration or session distribution? What's the best way to validate if sessions are being incorrectly handled in software instead of offloaded?
3. Are there known issues in PAN-OS 10.1.x related to memory or CPU management on the 3200 series that might manifest under certain conditions, even with low traffic?

I can provide sanitized outputs from relevant `show` commands. My next planned step is to engage TAC, but I prefer to arrive at that call with a thorough initial analysis.



   
Quote
(@devops_barbarian)
Honorable Member
Joined: 5 months ago
Posts: 439
 

You're focusing on the wrong dataplane stats. The mp and dp processes both spiking points to something in the session table itself, not the throughput.

Run `debug dataplane packet-diag set filter on` then `show session info`. Look for a huge number of sessions to a small set of IPs, especially on non-standard ports. I've seen stale SIP or broken SSL decryption sessions do this - low bandwidth but they churn the CPU.

Check your SSL/TLS service profiles too. A misconfigured TLS 1.3 policy can cause constant renegotiation that the box logs as normal sessions but murders the management plane.


Don't panic, have a rollback plan.


   
ReplyQuote
(@code_reviewer_anna)
Honorable Member
Joined: 5 months ago
Posts: 484
 

Great catch on checking the session table structure. Another thing that can cause similar mp/dp churn is a fragmented session cleanup. We had a case where a misbehaving client was creating thousands of short-lived sessions that timed out weirdly, causing constant table maintenance.

> Check your SSL/TLS service profiles

Absolutely second this. If you're doing decryption, add `ssl decryption mirror session` to your troubleshooting. It'll show you if the CPU is getting hammered by repeated handshake failures. Sometimes a server change on the other end can trigger this without any config change on your side.


Clean code is not an option, it's a sanity measure.


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

> verified session count is within specification

Is that based on total count or the session table distribution? The PA-3220 can handle the total session limit, but if 90% of those sessions are hitting the same rule or landing in a single virtual system, it can still thrash the CPU. The `show session distribution` command might show something weird the basic count doesn't.

Also, did the high CPU start after a specific software update? I saw a note on the support forums about a bug in one 10.1.x release causing extra dp overhead for certain app-id scans. A quick rollback to the previous version confirmed it for us.



   
ReplyQuote
(@crm_hopper)
Honorable Member
Joined: 7 months ago
Posts: 472
 

Session distribution is a good angle, but that table maintenance overhead usually shows up in flowd first. If mp and dp are both high, it's more likely a config being processed per-packet, not per-session.

The software update bug is a real rabbit hole. We chased that on a 10.1.6-h4 build for a week. TAC's final answer was "expected behavior" after they couldn't pin it down. Rollback worked for you because you got lucky.


CRM is a necessary evil


   
ReplyQuote
(@integration_maven)
Reputable Member
Joined: 6 months ago
Posts: 261
 

I've seen that exact pattern with configs processed per-packet. It's a different problem space than session table bloat.

If you're using QoS or a policy-based forwarding rule with a complex match, the dataplane has to evaluate it against every single packet, not just at session setup. That can drive both `dp` and `mp` up together because the management plane is constantly syncing the decision state.

The bug chase with 10.1.6-h4 sounds familiar - we hit a similar issue where a specific app-id group with too many members created per-packet overhead. It was technically "within spec" but performed terribly. Sometimes the workaround is to split a monolithic rule into two or three simpler ones.


IntegrationWizard


   
ReplyQuote
(@george7)
Honorable Member
Joined: 3 months ago
Posts: 572
 

That's a really good point about per-packet processing. A complex QoS rule could definitely be the culprit here.

Your mention of splitting a monolithic rule made me think of another scenario. Sometimes it's not just a complex match, but a rule with multiple services or applications that gets hit by traffic using many different ports. Each unique port/protocol combo can trigger a separate lookup in the policy set, multiplying the per-packet overhead.

You've probably seen it, but the `debug` output from `show running resource-monitor ingress-backlog` can sometimes expose these per-packet bottlenecks more clearly than just looking at CPU.


Keep it constructive.


   
ReplyQuote