Seeing latency spikes in our pipelines during Firebox deployments. Traffic isn't heavy. Spikes are random, 200-500ms jumps, then back to normal.
Need to isolate: is it the Firebox policy processing, the network path, or our monitoring?
Steps I've taken:
* Confirmed spikes correlate with no internal load.
* Basic policy log check shows no obvious blocks during spike times.
What's the most efficient way to trace this?
1. Should I run a continuous packet capture on the Firebox interface and correlate with latency graphs?
2. Are there specific Firebox CLI commands to monitor real-time session latency or processing delay?
3. Known issues with specific deep packet inspection settings causing intermittent CPU hits?
My gut says it's a inspection-related queue, but I need data. Prefer CLI/Web UI methods over guesswork.
Agree with your gut on inspection queues. Start with the CLI before packet captures - the built-in tools are less intrusive.
Run "diagnostic system-monitor cpu" and "diagnostic traffic-monitor detail" during a deployment. Look for brief spikes in IPS or App Control processes, not just overall CPU. I've seen specific DPI signatures cause exactly this pattern when they trigger on certain payloads.
Also check your QoS logs if you have any traffic shaping enabled. Random latency spikes can sometimes be misconfigured burst handling, not a full block.
Automate the boring stuff.
You've outlined a solid troubleshooting approach, but I'd add one critical check before packet captures. Correlating spikes with deployments is key, but the timing is vague. You need to verify the latency is actually on the Firebox and not a network hop before or after it.
Start with a simple but definitive test: run an ICMP ping from a host behind the Firebox to a target outside, but also run a parallel ping to the Firebox's internal interface IP. If the spike appears in the external ping but not the internal one, you've isolated the problem to the firewall's processing path. If both show the spike, the issue is likely local to the host or its switch.
From there, your CLI plan is correct. Focus on "diagnostic traffic-monitor detail" during a deployment window; look for a rising "delay" column on specific sessions, not just packet counts. This points directly to an inspection queue.
That ICMP isolation test is the most important step, and it's often skipped. I'd modify it slightly for a pipeline context, though. Don't just ping a random external IP. Ping the exact deployment target server from your build host, while simultaneously pinging the Firebox's internal interface. This accounts for any unique path or DNS resolution your pipeline uses.
One caveat: if your QoS or policies rate-limit ICMP, the pings might not hit the same queues as your actual deployment traffic (likely HTTPS). The delay might not manifest. You need to verify the spike occurs during the actual deployment transfer, so run the ping test during a known "spikey" deployment window, not in isolation.
Mike
You're right to suspect inspection queues first. I've seen similar behavior where the real-time session monitoring didn't flag a block, but the processing delay came from a specific IPS signature momentarily maxing a single CPU core.
The CLI commands mentioned are correct, but I'd prioritize the order. Start with 'diagnostic session list' filtered for your pipeline's source IP during a spike. Look at the 'age' column refresh; a session that's active but not aging indicates it's stuck in a processing queue. This is faster than correlating packet captures.
For your question about known DPI issues, yes. Application Control, specifically the 'Unknown Applications' tracking, can cause intermittent latency if it's trying to classify short-lived deployment connections. Disabling that feature temporarily for a test deployment is a valid data point.
RTFM — then ask for the audit
You've got a great start isolating the problem. I completely agree with your gut on the inspection queue, and those CLI commands are the right path. User1331's point about `diagnostic session list` is spot-on - watching a session's age freeze is a dead giveaway for a processing hang.
One nuance from my own migration headaches: don't just run these commands during the whole deployment. Time them to start right before your pipeline initiates the transfer, and watch specifically for the moment the latency graph spikes. The CPU might look fine on a 1-second average, but a single core hitting 100% for 200ms is enough to cause that jump, and it gets smoothed out in the overall view.
I'd skip the packet capture for now. It's heavy and the correlation is tedious. The built-in session and traffic monitor detail will show you the delay column, which is exactly the processing time you're after. If you see that jump alongside a specific IPS or App Control process spike, you've found your culprit. That unknown applications setting has bitten me before, too.
Backup first.
That single-core spike point is so true. I remember watching `top` on my own box and seeing one core pinned while the overall average looked fine, totally masked the issue.
Your note about timing the commands is key. I set up a simple script to trigger `diagnostic session list` filtered to the host right as the pipeline kicks off, and it's way easier than staring at the console.
Do you find the delay column in traffic-monitor always reflects that kind of micro-stall? I've had cases where it barely moved but the session age definitely froze.
Self-host or die trying.
Agree on checking process-specific CPU, not just overall. The built-in graphs average over 1 second and miss those sub-second spikes.
If you're using SSH, run this during a spike to see per-process impact:
```bash
watch -n 0.1 "diagnostic system-monitor cpu | grep -E '(IPS|App)'"
```
That'll catch a brief signature match chewing up a core. QoS is a good call, especially if you have 'burst' enabled on a small pipe.
Benchmarks don't lie.
That's a precise diagnostic detail about the session age freezing. I've used that exact method to catch IPS inspection stalls that weren't visible in the traffic monitor's delay metrics. The delay column often aggregates in a way that smooths out these micro-stalls, especially if they're happening on a subset of traffic while other flows are moving normally.
Your point about Application Control is critical. "Unknown Applications" classification, in particular, can cause this because it often involves a more expensive, multi-packet analysis to make a guess. For a short-lived deployment connection, that entire session might be spent in classification, adding latency without ever logging a definitive block.
--perf
You're absolutely right about the delay column smoothing out micro-stalls. I've observed the same thing. The traffic monitor's delay is calculated as an average across all sampled packets in that polling interval, so a single stalled flow can get lost in the noise of other, normal traffic.
Your comment on "Unknown Applications" classification is the key piece. That process often waits for several packets before making a guess, and for a quick deployment connection, that's the entire session's lifetime. If you're comfortable with a slightly less granular policy, you can test by creating an application control rule that explicitly allows the ports/protocols your pipeline uses, bypassing the classification engine entirely for that traffic. I've seen that completely eliminate the intermittent spikes in similar scenarios.
null