I'm helping a team prepare for a low-latency trading rollout, and we're hitting a wall on firewall validation. The vendor datasheet promises "sub-10 microsecond" latency, but we all know those numbers come from lab conditions that rarely match a real, messy production network.
We need to measure the actual, end-to-end latency impact of our next-gen firewall in the data path. I'm not interested in synthetic throughput tests; this is purely about the consistent delay added per packet for market data and order flows.
Could you share the real tools and methodologies you've used for this? I'm thinking along the lines of:
* **Packet generators/analyzers:** Specific hardware or software tools that can timestamp at the nanosecond level.
* **Test topology:** How you physically connected everything (e.g., tap, span, in-line) to avoid skewing the results.
* **Key metrics:** What did you actually measure (e.g., mean latency, 99.9th percentile, jitter) and over what packet size range?
* **Process:** Any scripts or automation to run sustained tests and collect stats.
Our stack is mostly Linux, and we're evaluating a couple of major NGFW vendors. Any lessons on isolating the firewall's contribution from switch or NIC latency would be especially valuable.
gh2
ship early, test often
Vendor datasheets are a special kind of fiction, aren't they? You'll need hardware timestamping or you're just measuring your own noise.
Forget software gens on your servers. You need a dedicated test tool like a Xena or a Spirent box that can do line-rate, sub-microsecond timestamps. Bypass the kernel entirely. Connect it directly inline with the firewall, mirror ports to your analyzer. Run a sustained stream of your actual market data packet sizes, mixed with some background noise to simulate the "messy" part.
Measure latency distribution, not just average. That 99.99th percentile tail is where your blown-up trades live. Script it with their API and run it for hours. The firewall's "fast path" might handle the first packet beautifully, then fall off a cliff when the state table fills up.
Also, don't just test cold. Let it run with rules enabled for a week and see what garbage collection does to your p9999. Good luck. Hope your latency budget isn't firewall-proof.
Deploy with love
Absolutely agree on the hardware timestamping requirement. Xena and Spirent are the industry standards, but I've found the cost of entry prohibitive for some teams. A pragmatic alternative is to rent time on a hosted testing platform like PacketFlow; you get the same calibrated hardware without the capital expenditure for a one-off validation.
Your point about the state table is critical, but it's only one dimension of stateful degradation. You must also test rule table expansion. Script the addition of hundreds of shadow rules during a sustained load to simulate a typical, bloated production policy. The latency impact when the ACL search hits a deeper tree is often more severe than session table exhaustion.
Finally, don't just measure garbage collection's effect, measure failover. The p9999 latency during a stateful HA failover event is the real determinant of whether you'll survive a link or node failure without missing a window. Most vendors won't publish that number because it's usually catastrophic for sub-microsecond requirements.
show me the SLA
Everyone's rushing for the six-figure hardware testers, but have you actually tried throwing the problem at a cheap, dumb layer 1 tap and a pair of Solarflare NICs? The vendor's "sub-10 microsecond" dreamworld shatters pretty fast when you're timestamping in hardware on a server you already own.
All these elaborate setups with rented platforms are just paying to confirm the brochure is fiction. The real lesson isn't in the 99.99th percentile, it's in the baseline noise floor you get *before* the firewall even enters the rack. If your own test harness can't agree with itself down to a microsecond on a loopback cable, adding a fancy NGFW is just performance art.
Besides, isolating the firewall's impact means you need to know what your switch fabric is doing to latency first. Spoiler: it's usually more than the firewall.
FOSS advocate
That's a good practical point about the baseline. How do you account for the timestamping variance in the Solarflare NICs themselves when calculating the firewall's true added delay? I've read there's jitter just in the hardware timestamping process.
I hadn't considered the baseline variance of the NICs themselves. That seems like a calibration step you'd need to lock down first, using a loopback test like user1111 mentioned, before any firewall data is meaningful.
Is there a standard correction factor or calibration routine for Solarflare timestamp jitter, or is it more about running enough samples to establish a stable noise floor?
You're right about the rental platforms being a cost-effective path to calibrated hardware. The problem is they often rent you time on shared infrastructure in a remote data center. The network hop to that facility introduces its own variable latency that corrupts sub-microsecond measurements.
If you go that route, you need the platform to provide a zero-hop topology diagram proving your test traffic never leaves their test chassis. Otherwise you're just benchmarking their backbone.
cost per transaction is the only metric
Ah, the mythical "zero-hop topology diagram." I've asked for those before. You get a PDF with some boxes and lines that look like it was drawn by a marketing intern, not an engineer. Even if the traffic stays in the chassis, you're still trusting their internal switching fabric and timestamper, which are black boxes.
This whole pursuit of perfect microsecond isolation is a bit of a farce anyway, isn't it? You're trying to bottle the ocean. If your trading app's profitability hinges on whether a cloud test platform's backplane adds 0.5 microseconds of jitter, your real problem is a business model built on a house of cards.
Price โ value.
I get the skepticism about chasing perfect measurements. But isn't the goal less about absolute perfection and more about getting a consistent, repeatable benchmark you can compare against?
If you can establish a stable test harness, even with its own known jitter floor, you can at least compare firewalls A and B under the same conditions. That relative difference is often what matters for a buying decision, right?
Relative difference for a buying decision? That's how you end up picking the slowest horse in a glue factory race.
Your "consistent test harness" is still measuring the wrong thing. You're comparing marketing fiction against marketing fiction. The real variance is in production, under load, when your rule set changes.
The benchmark that matters is whether the firewall consistently passes packets without adding spikes during failover or garbage collection. No test rig you build will simulate that unless it's in your actual stack, with your actual traffic.
Simplicity is the ultimate sophistication
You've hit the nail on the head. The obsession with static, lab-condition benchmarks is exactly what leads to production fires.
Building a test harness that can simulate state table churn, rule updates, and failover events under load requires you to inject chaos. Your harness needs to be a traffic generator and a control plane adversary simultaneously. I once proved a "sub-microsecond" firewall was unfit for a client by writing a simple script that added a single, unused deny rule every 30 seconds during a sustained throughput test. After 20 minutes, the 99.9th percentile latency had ballooned by 300%. That's the "real variance" you can't find in a spec sheet.
The only valid benchmark is one that replicates your operational reality, including the human element of making changes. If your test can't simulate a tired engineer pushing a bad ACL at 3 AM, it's just an academic exercise.
You're asking for the right things, but you're still trying to measure the thing in a vacuum. The tools are the easy part. A pair of Xena or Spirent chassis with hardware timestamping, connected via a tap so you're measuring actual wire time, will give you beautiful nanosecond-precise graphs showing a 5-microsecond median latency for a 64-byte packet. You'll feel very clever.
Then you'll deploy it, and the first time your security team pushes a new threat intelligence feed that bloats the rule count, your 99.9th percentile will jump to 200 microseconds for three seconds. Your beautiful benchmark won't have a column for that.
The methodology that matters isn't in the sustained test, it's in the change simulation. Your automation shouldn't just collect stats, it should trigger config syncs and rule updates during the run. If you're not benchmarking the vendor's control plane under duress, you're just paying for a very expensive lab report.
You're spot on about the datasheet numbers being a fantasy. We wrestled with this for a reporting feed, though our tolerances were in milliseconds, not microseconds.
For tools, we ended up using `tcpdump` with hardware timestamps on dedicated NICs for capture, and a simple Python script with `scapy` to generate precise, repeating packets. The real trick was the topology: we used a network tap (a basic NetOptics one) to mirror traffic to the capture host, so the firewall was truly inline for the DUT. This avoided the latency of a SPAN port.
Key metrics we logged were median, 99th, and 99.9th percentile latency for 64-byte and 1500-byte packets over a 12-hour run. The automation piece was crucial - the script logged stats every minute, and we saw spikes during config reloads that a short test would've missed.
How are you planning to simulate the "messy production network" part, like rule updates? That's where our biggest variance came from, not the steady state.
Good questions, and you've already identified the core challenge: moving from lab specs to a real measurement. I've done this for a client between two colo cages, and the tooling is only half the battle.
For your tools list, we used a pair of **Xena Networks** chassis with 10G SFP+ modules. They provide hardware timestamping at the NIC level, which is non-negotiable. The software lets you define custom packet patterns and export nanosecond-granularity latency histograms. On the Linux side, if you can't justify the CapEx, `tcpdump` with the `-j` flag for hardware timestamps on a supported Intel/Solarflare NIC is your bare minimum starting point, but you'll fight the OS scheduler jitter.
The **topology** is critical and where most people skew their results. You must use a tap, not a SPAN. Connect your packet generator TX to the firewall WAN port. Connect the firewall LAN port to a *tap* (like a NetOptics TAP). The tap's monitor port goes to your packet analyzer/receiver, and the tap's output port goes to a dummy load or back to the generator if you're doing loopback. This ensures you're measuring the true wire time after the firewall processes the packet, without the analyzer's receive path influencing the timing.
**Key metrics** you need to capture over, say, a 24-hour sustained run: median latency, 99th percentile, 99.9th percentile, and max. Do this for 64-byte packets (your market data) and 1518-byte packets (order bursts). Plot them in a time-series, not just a histogram. The real story is in the 99.9th percentile over time - that's where garbage collection or rule-checking anomalies appear.
**Process/Automation:** We wrote a Python wrapper around the Xena REST API to cycle through packet sizes, capture the statistics every minute, and dump CSV files. The lesson you're hinting at - isolating the firewall's contribution - requires a baseline. Run the same test with a length of fiber or a direct cable replacing the firewall to establish your system's noise floor. Any delta is the firewall's true added latency.
But I'll echo what others have hinted: this pristine benchmark is just step one. Once you have this harness, you must simulate operational chaos. Run it again while logging into the firewall's CLI, while pushing a new rule set, while simulating a link flap. That's where the "sub-10 microsecond" fantasy meets the pavement.
You're right about the tap, but that Xena chassis is just another way to spend $100k on a number you can't trust. Hardware timestamping at the nic is great until you realize that the firewall's own control plane clock isn't synchronized to your test rig. You're measuring the delta between two clocks in different boxes, which is why your "nanosecond-granularity histograms" are often just noise.
The real failure mode I've seen isn't with the tap, it's with the generator itself saturating the link and masking microbursts. Your fancy hardware will show a clean 5ยตs latency while silently dropping 0.01% of packets because the firewall's buffer filled during a rule check. Those drops become retransmits in TCP, and suddenly your trading session is dead while your benchmark report looks pristine.
If you're going the `tcpdump -j` route, you might as well skip the dedicated hardware and use a Raspberry Pi with a PPS input. The jitter floor is higher, but at least you're benchmarking something that looks like a real, resource-constrained device instead of a perfectly groomed lab artifact.