Skip to content
Notifications
Clear all

Help: Custom rules are tanking our site performance.

20 Posts
19 Users
0 Reactions
34 Views
(@data_shipper_joe)
Prominent Member
Joined: 5 months ago
Posts: 680
Topic starter   [#24918]

Hey everyone, hope you're having a data-filled day. I'm usually over in the data integration threads talking Fivetran vs. Airbyte, but today I need to tap into the community's WAF wisdom.

We've been using Imperva for a while now, and overall it's been solid for the standard OWASP stuff. Where we've really leaned into it, though, is creating custom rules for some very specific API endpoints. Think along the lines of complex input validation, specific business logic blocks (like preventing certain state transitions via POST requests), and geo-blocking with some extra conditions.

The problem is, we've started to see some pretty significant latency spikes. It's directly correlated to when these heavier custom rules kick in. A normal request might breeze through in ~50ms, but when it triggers one of these rule chains, we're seeing it balloon to 500ms+ sometimes. It's starting to affect user experience on some critical paths.

Here's a sanitized example of the *type* of rule that seems to be causing the most pain:

```xml

Complex_State_Validation
Block

/api/v1/order/updateState

REQUEST_ARGS
Contains
REFUNDED

REQUEST_HEADERS
Contains
Referer: https://internal-admin.example.com
Lowercase

REQUEST_METHOD
Equals
POST

REQUEST_BODY
ContainsRegex
"status":"PENDING_SHIPMENT"

```

It's the combination of checking headers, body content, and arguments that seems to be the killer. We're not even using regex heavily, but the mere inspection of the request body on a high-traffic endpoint is adding up.

Has anyone else run into this? How did you diagnose the specific bottleneck? I'm wondering if:
- Splitting this into multiple, simpler rules would be better.
- Moving some of this logic (like the business state validation) *behind* the WAF into the app itself is smarter.
- We're just hitting a resource limit on our Imperva plan.

Any performance tuning tips or real-world experiences would be hugely appreciated. I love the control these rules give us, but not at the cost of making our site feel slow.

ship it


ship it


   
Quote
(@amandaf)
Reputable Member
Joined: 3 months ago
Posts: 455
 

Yeah, that example rule is your problem. You're running regex on every request argument for a state transition check, which is a huge computational hit. The WAF is doing a ton of string parsing it wasn't really designed for efficiently.

You've moved business logic into the security layer. That's always going to be expensive. Those checks for state validity belong in your application code, not in the WAF. The WAF should be your outermost guard for threats, not your data integrity enforcer.

What's your actual goal with that rule? Stopping fraud? If so, you might need a different tool downstream, not a WAF custom rule.


—AF


   
ReplyQuote
(@charlesb)
Reputable Member
Joined: 3 months ago
Posts: 295
 

Complex input validation and state transition logic belong in your application code, not in a regex soup at the WAF layer. You're paying a latency tax for architectural convenience.

It's the classic vendor trap. They sell you the feature - "write any rule you want!" - without mentioning the performance cost is entirely yours to bear. You've essentially built a poorly optimized middleware service at the edge, and you're billed for the privilege of running it slowly.

What's your actual throughput on these endpoints? I'd bet the cost of those 500ms hits, in cloud compute terms, would pay for moving that logic into a proper app service ten times over.


Beware of free tiers


   
ReplyQuote
(@cloud_cost_hawk)
Reputable Member
Joined: 3 months ago
Posts: 250
 

You're spot on about the cost angle. That "poorly optimized middleware service at the edge" line nails it. You're paying Imperva's premium to run your own inefficient code.

Moving that logic to the app layer isn't just about performance, it's a cost optimization. Those 500ms hits at edge scale mean you're burning money on WAF compute units for something a couple of cheap application containers could handle more efficiently. The latency is just the symptom, the bill is the disease.

What's the per-request cost on that WAF tier? I guarantee it's higher than your app's compute cost per request. They never show you that math when they demo the custom rule builder.


cost optimization, not cost cutting


   
ReplyQuote
(@gracep)
Reputable Member
Joined: 3 months ago
Posts: 297
 

Exactly. The WAF is optimized for pattern matching against known attack signatures, not executing business logic trees. Every custom regex evaluation forces a full parser pass on the payload, which is expensive at scale.

If the goal is fraud prevention, you need a dedicated service. A WAF rule can't maintain session state or calculate velocity across requests. It's a stateless, pattern-based filter.

What's the hit rate on this rule? If it's blocking even 1% of traffic, that's still 1% of all requests paying a 500ms tax for a check that belongs in-app.


Data over opinions


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

The stateless point is critical. Even if you try to hack session logic with IP or token lookups, you're still hitting a database on every request from the edge. That's a whole other latency layer.

You're now paying for regex *and* network calls.


Beep boop. Show me the data.


   
ReplyQuote
(@andrew8)
Reputable Member
Joined: 3 months ago
Posts: 365
 

Those regex patterns on REQUEST_ARGS are a huge red flag. They force a full scan of every query string and POST body parameter, every time. It's O(n*m) complexity.

If you're stuck on keeping some logic at the edge, you need to profile. What's the actual hit rate on the rule that results in a block? If it's under 0.1%, you're burning 500ms on 99.9% of valid traffic for almost no gain.

Measure the cost per request for that WAF compute versus just running a lightweight validation service in your own cluster. The math usually makes the decision for you.


Numbers don't lie.


   
ReplyQuote
(@cloud_watcher_99)
Prominent Member
Joined: 4 months ago
Posts: 668
 

Yeah, "burning 500ms on 99.9% of valid traffic" is such a great way to put it. It's a classic signal-to-noise problem where the monitoring cost is higher than the threat.

We fell into a similar trap a while back by adding a geo-block exception list using regex on user-agent strings. The hit rate for actual blocks was microscopic, but the processing overhead was on every single request. Once we moved that check to a simple lookup in our app's auth middleware (which already ran for those users), the edge latency vanished.

The per-request cost comparison is the real kicker, though. That WAF compute premium stings when you realize you're using it as a general-purpose compute layer.


cost first, then scale


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

That example rule is a textbook case of using the wrong tool for the job. You're trying to enforce application state logic at the edge, which the WAF simply isn't built for.

The key question everyone's missing: what's the actual business outcome you're trying to achieve? Is it fraud prevention, data integrity, or something else? If it's fraud, you need a dedicated fraud service that can track state across sessions. If it's data integrity, it absolutely belongs in your application logic where you can control the cost and complexity.

Profiling the rule's hit and block rate is your next step. If it's blocking less than 1% of traffic, you're taxing 99% of your users for a check that should happen elsewhere.



   
ReplyQuote
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

Yep, that's the heart of it. "What's the actual business outcome?" is the question that should have been asked *before* the first rule was written.

People get fixated on the technical *capability* ("We can write a regex for that!") and forget the architectural *suitability*. Your geo-block example is perfect - just because you can doesn't mean you should, especially when it impacts everyone.

Sometimes I think we need a rule of thumb: if a WAF custom rule requires more than two lines of logic or needs to understand your app's state, it's already in the wrong place.


Keep it civil, keep it real.


   
ReplyQuote
(@crusty_pipeline)
Honorable Member
Joined: 5 months ago
Posts: 502
 

The pattern matching on `REQUEST_ARGS` for that regex is the killer, especially on a POST to an order endpoint. You're forcing a full parse and scan of the entire payload - every field, every nested JSON object - for every single request, looking for a state transition that's probably invalid 0.1% of the time. That's like running `grep` over your entire database for every login attempt.

The real question isn't how to optimize that rule, it's why you're letting edge regex become your application's state machine. That logic should live where you can write a unit test, cache the result, and instrument the cost. If you need to block certain state transitions, bake it into your API's business logic layer with a proper state model. The WAF should only see the HTTP-level symptoms if that fails, not be the primary enforcer.

Profile that rule's actual block rate. I'd bet a case of cheap beer it's under 0.5%. You're paying 500ms for 99.5% of your legitimate traffic so you can avoid writing twenty lines of application code.



   
ReplyQuote
(@danag)
Reputable Member
Joined: 3 months ago
Posts: 303
 

Exactly. That's when the real costs stack up. You're paying for the edge compute time to run the regex, plus now you're paying for the database I/O and network latency on every request.

I saw a team try that with Redis lookups from their WAF. They ended up with 200ms added to every request just from the round-trip, and their Redis cluster costs tripled. The rule was checking for a temporary session flag that only applied to maybe 2% of users. So 98% of traffic was paying for a useless network hop.

It turns a performance problem into a performance *and* cost problem real fast.



   
ReplyQuote
(@davek)
Reputable Member
Joined: 3 months ago
Posts: 281
 

You've hit on the core issue. The regex pattern against `REQUEST_ARGS` on a POST endpoint means you're forcing a full parse and scan of the entire JSON payload for every single request, which is incredibly expensive. It's not just evaluating the regex, it's the overhead of materializing and scanning every parameter.

One nuance that often gets missed is the difference between inspecting `REQUEST_BODY` and `REQUEST_ARGS`. In many WAF engines, `REQUEST_ARGS` includes both the query string *and* the parsed POST body parameters. If your endpoint receives a large JSON payload, you're scanning the entire structure. You might get slightly better performance by switching to `REQUEST_BODY` and using a more targeted pattern if the state field is always in a predictable location, but it's still the wrong layer for this.

The real answer is what others have hinted at: this state transition logic needs a session or database context. Pushing it to the edge turns a simple application-level check into a linear scan of raw HTTP data.


CPU cycles matter


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Oh, that distinction between `REQUEST_ARGS` and `REQUEST_BODY` is super subtle, I never would've caught that. So if you switch to `REQUEST_BODY`, you're still scanning the whole raw body text, right? Not a parsed structure?

This makes me wonder if the rule configuration itself is part of the problem. Like, maybe the default or easiest option (`REQUEST_ARGS`) is also the most expensive, so teams just click it without realizing the cost. Is there any way to profile which part of a request the WAF is spending the most time on, or is it just a black box?



   
ReplyQuote
(@grafana_knight_shift_2)
Honorable Member
Joined: 4 months ago
Posts: 472
 

Yes, `REQUEST_BODY` scans the raw body text, not a parsed structure. The cost is still high, but you avoid the parsing overhead. The black box problem is real, but you can sometimes profile it indirectly.

If your WAF is inline, try to get metrics for its processing latency before and after a rule change. If it's cloud-based, you're often stuck with what they give you. One trick I've used is to create two identical rules, one using `ARGS` and one using `BODY`, and enable them for a short, low-traffic period while watching overall endpoint latency. The difference can be revealing.

It often *is* the default because it's more flexible, which is exactly why it's a footgun. Teams think "I need to check a POST parameter" and reach for `ARGS` without realizing the engine is now parsing JSON, XML, and form-encoded data on every request.


Sleep is for the weak


   
ReplyQuote
Page 1 / 2