We've been running Radware's application delivery controllers in a primary/backup pair for our east-west service mesh ingress points, specifically terminating mTLS from Istio gateways before traffic hits our core application tier. Our deployment is fully automated via Terraform and Ansible, with firmware upgrades handled through a controlled pipeline. The previous firmware version (we were on `Release_20.1.00`) was stable for our workload patterns.
This week, we proceeded with a planned upgrade to `Release_20.2.01` following the vendor's recommended staged rollout. Post-upgrade, our monitoring (Prometheus/Grafana dashboards sourcing metrics from the ADCs and supported by application-level tracing) showed an immediate and sustained **40% increase in p99 latency** for requests passing through the devices. The increase was isolated to the newly upgraded primary unit; traffic failing over to the backup on the older firmware immediately returned to baseline latency. The spike was consistent across all services, pointing to a systemic issue rather than a specific routing or policy change on our end.
Key observations from our diagnostics:
* CPU utilization on the ADC remained within normal bounds (<50%).
* Connection table counts and memory usage were unchanged from pre-upgrade levels.
* The latency introduced appeared to be in the SSL/TLS processing path, even with hardware SSL acceleration enabled. Simple HTTP traffic saw a less pronounced, but still noticeable, increase.
* No errors or warnings in the device logs correlated with the slowdown. All health checks passed.
We attempted the following mitigations without success:
* Verified and reapplied our optimal cipher suite configurations.
* Disabled and re-enabled advanced features like HTTP/2 multiplexing and compression.
* Conducted a packet capture, which showed increased delta between TCP ACK and the first application byte.
Our rollback procedure to `Release_20.1.00` was executed, and latency metrics normalized immediately upon completion. The rollback itself was straightforward, but the incident has disrupted our confidence in the automated upgrade path.
Has anyone else encountered performance regressions, particularly in TLS handshake or request processing times, with the latest firmware? We are particularly interested in experiences from environments using the ADCs in front of a Kubernetes or service mesh layer, as the request profile (many small, encrypted HTTP/2 streams) might be a factor. We are now compelled to build significantly more exhaustive performance testing into our firmware evaluation pipeline, but I'm curious if there are specific known issues or configuration adjustments required post-upgrade that we may have missed.
Ouch, latency spikes after a firmware update are the worst kind of predictable surprise. Your diagnostic approach is solid, isolating the variable to the specific upgraded unit.
You mention CPU remained within normal bounds, but did you check memory allocation or SSL session cache behavior in the new version? Sometimes the "optimizations" shift the bottleneck to a different subsystem. I've seen a vendor patch that quietly halved the default TLS ticket cache, causing constant renegotiations and p99 hell.
Good call rolling back. Besides opening a high-severity ticket with Radware, you might want to check their community forums - someone else has probably hit this and posted a workaround config tweak. Saved my team a three-week support loop once.
- elle
The TLS cache angle is plausible. In my past incident involving a different vendor, the new firmware's debug logging for the SSL module was enabled by default at a verbose level, causing excessive small writes and filesystem contention. It didn't show in overall CPU but spiked iowait.
Radware's release notes sometimes bury those tunable defaults in an appendix. The session cache size and SSL log level would be my first two checks after the config parity.
Absolutely, the debug logging default change is a classic silent killer. I've run into similar with an HAProxy upgrade where the `debug` verbosity level for a seemingly unrelated module started logging every DNS resolution to disk. The latency didn't come from the log writes themselves, but from the mutex contention on the shared log buffer across all worker processes. The p99 would balloon while average latency and CPU looked fine.
You're right to point out the release notes appendix. I've found that the actual 'Changes in Default Behavior' are often listed separately from the main bug fixes, sometimes in a PDF appendix. It's worth a `diff` of the running config's non-explicit defaults between firmware versions, if the CLI allows it. The SSL log level might not even appear in your saved config if you never manually set it.
throughput first
Did you get a quote for extended support on the old version before rolling back? Some vendors will try to charge you for staying on a stable release past its standard term. It's worth checking your contract's maintenance clause.
That's a practical and often overlooked point. The financial operations angle on incident response is real. In our last vendor escalation over a similar regression, staying on the patched-but-stable older release required a specific waiver that did trigger a review of our support tier. The cost wasn't prohibitive for the quarter we needed it, but the procurement cycle to approve it added more delay than the technical rollback.
I'd suggest checking not just the maintenance clause, but also the software lifecycle document for that product line. Radware typically lists "End of Engineering" and "End of Support" dates per major release train. If 20.1 is still within the engineering support window, you're likely covered without extra fees. The pressure to upgrade usually comes when you're nearing the end of that period and need security patches.
Your contract might also have language about "commercially reasonable efforts" to support prior versions during a proven regression, which gives some leverage.
Latency is a liability
Good point on the diff, but you can't always trust the CLI to show changed defaults. Some vendors keep those in a separate binary config that doesn't diff cleanly. Been burned by that before.
The HAProxy mutex issue is spot on. I've seen the same pattern with syslog over TCP as a default after an upgrade. Suddenly every log line is a blocking network call.
Check if they changed the default audit or telemetry level too. That's another silent performance tax.
Trust but verify.
The isolation to the upgraded unit is the critical signal. Since you've ruled out your own config changes via automation, the next step is a granular resource check beyond aggregate CPU. My comparison spreadsheet for similar incidents tracks three culprits that don't show in overall utilization:
* **Interrupt balancing or RSS queue changes** between firmware versions, which can pin a single core.
* **Memory allocator fragmentation** from a new default, increasing latency for new connections.
* **Kernel scheduler tweaks** affecting poll() or select() loops in the TLS handshake path.
Have you compared `perf` output or equivalent internal latency histograms between the two units during the same traffic pattern? That usually pinpoints the subsystem shift.
Measure twice, buy once.
`perf` is a good call, but pulling internal histograms from a black-box appliance is often impossible. You need the vendor's cooperation.
The RSS queue point is key. I've seen firmware change the default number of receive queues without updating the driver's CPU affinity map. One core hits 100% softirq while the overall CPU metric looks fine. That matches a latency spike with normal aggregate utilization.
You should check the interrupt distribution per core on both units during load. If the new firmware switched to a different NIC driver or queue scaling algorithm, that's your smoking gun.
cost per transaction is the only metric