We started with similar team-based tagging but found it lacked the granularity needed for effective cost allocation. Our current approach uses a multi-dimensional tagging system where each rule gets at least two tags: one for the owning team (like `team:security`) and one for the protected asset or endpoint (like `asset:checkout-api`).
The real insight came when we added a third tag for the primary threat category, like `threat:scraping` or `threat:credential-stuffing`. This let us pivot the cost data in three ways. We could see not just which team's rules were expensive, but which specific business functions bore that cost and what type of attack was driving it. For example, we identified that `threat:scraping` on `asset:product-catalog` was our single largest cost center, which led to a more targeted optimization than just turning rules on or off.
The overhead to maintain this is minimal if you templatize rule creation. The key is ensuring the asset and threat tags align with your existing internal classifications.
Data over dogma
Tagging rules for cost attribution is a logical first step, but your weekly report pull raises a more foundational question: what's your empirical basis for determining which paths are "non-critical"? Without a quantitative model tying security posture to business risk, you're still operating on gut feeling.
Your script should be validating that assumption. For each endpoint tagged as non-critical, correlate its traffic volume, the percentage of mitigated requests, and the estimated business impact if a blocked legitimate request occurred. You might find that some high-traffic, low-risk endpoints are where your cost inefficiency is concentrated, not just the rules themselves.
The deeper issue with vendor APIs is that their "mitigated traffic" metric is often a black box. Have you attempted to reconcile their counts with your own application logs to verify the volume they're charging you for? Discrepancies there can be a significant source of wasted spend.
Trust but verify.
You're right to push for empiricism over gut feel. That "estimated business impact" is the hard part. Most teams get stuck there and the model falls apart.
We started by forcing product owners to assign a simple risk score - low, medium, high - to each endpoint for confidentiality, integrity, and availability. The screams were loud, but it gave us a defensible baseline. Without that, your "non-critical" tag is just an opinion.
On the black box billing, yes, always reconcile. We found a persistent 3-5% overcount on mitigated requests versus our own WAF logs. The vendor's explanation was "processing delays" on their side. That discrepancy alone paid for the logging pipeline to track it.
Trust but verify – and audit
Oh, forcing product owners to score risk is a bold move. We tried that and it turned into a months-long debate about what "integrity" even means for a read-only API. 😅
Your point about logging to verify the vendor's metrics is 100% on the money. We built a small pipeline that samples and matches requests between our WAF and their billing API. The delta wasn't huge, but having the data to push back on a bill felt priceless. Did you ever manage to get a credit for that 3-5% overcount?
Dashboards or it didn't happen.
Integrating the Imperva API for our weekly pulls was mostly straightforward, but you have to be careful with the pagination. Their documentation says to use the 'offset' parameter, but we found the 'from' and 'to' timestamps for date ranges needed very precise formatting to avoid missing data at the boundaries, which skewed our weekly totals at first.
The data formatting itself wasn't a fight, but you have to plan for the nested JSON structure of some reports, like the security events feed. We wrote a small parser that flattens the key fields we need, such as rule ID and client IP, into our log schema. My main caveat is to watch for silent changes in field names between API versions. We had a field called "action" suddenly split into "mitigation_action" and "policy_action" without any warning in the release notes.
Are you planning to use a specific language or framework for your monitoring stack? I'm curious if others have hit similar issues with the timestamp granularity in their responses.
Your approach of weekly report pulls is the right start, but I'd question the "simple Python script" characterization. Without a proper statistical baseline for traffic patterns, you're just tracking totals, not diagnosing anomalies.
The real inefficiency often isn't the rules you have in block mode, but the volume of requests being evaluated against those rules in the first place. A better integration would feed your weekly Imperva data into a time-series model that flags weeks where the *ratio* of mitigated to total traffic deviates significantly from the historical norm. That's your signal to investigate, not just a raw cost increase which is expected with growth.
Also, your script should validate the vendor's "mitigated" classification. We found a 7% discrepancy between what Imperva billed as mitigated and what our internal logs showed as actual actionable threats after filtering out false positives from known-good bots. Are you reconciling those counts, or just taking the billing data as fact?
p-value < 0.05 or bust
Your weekly report pull is a necessary step, but you need to validate the source data before feeding it to your models. The Imperva API's definition of a "mitigated request" isn't always what you'd assume. We audited the raw logs and found a substantial portion were for requests blocked at the TCP layer before they even reached our application or WAF rules, yet they still counted toward the billing metric. This is traffic you cannot meaningfully tune via rule modes.
Before you build any anomaly detection on the weekly aggregates, I'd verify that your script is filtering out these network-layer mitigations if they're irrelevant to your security posture review. Otherwise, your analysis of rule efficiency is based on a flawed denominator.
Exactly. If you're not peeling back the billing metric to see what's inside, you're just budgeting for their product's inefficiency. We found the same split - a chunk of our "mitigated" volume was just TCP resets from their DDoS mitigation, which we couldn't tune and didn't align with our security review cycles.
Our workaround was to tag those network-layer actions in our log parser and exclude them from the "tunable" cost pool. It made the remaining rule-based costs actually actionable. Did your audit show if that TCP-layer percentage was stable, or did it spike with traffic growth? Ours seems to drift, which makes static filtering risky.
Data over dogma.
That "set-and-forget" feeling is exactly what gets you. The bill doubling isn't just about more traffic, it's about the ratio of mitigated to total traffic creeping up.
Your weekly report pull is a good move, but you need to split that "mitigated" volume immediately. A big chunk is probably legitimate bot traffic hitting rules in block mode, like you found, but a surprising amount can be network-layer DDoS mitigation you can't tune with rule modes.
I'd make your script's first step categorizing the mitigation actions. Separate the TCP resets from the actual WAF rule triggers. That gives you a true cost for the security posture you can actually control versus a tax for just being online.
You've hit on the critical first step with tagging rules for review. That weekly Python script will be your lifeline, but its effectiveness hinges entirely on how you define "mitigated traffic" before it even gets to your analysis.
As others have noted, the vendor's billing metric is a composite. When your script pulls the weekly data, you must immediately dissect it. A significant and often growing portion of that "mitigated" volume is likely TCP resets from the DDoS layer, not WAF rule actions. These are not tunable via rule modes. If your script doesn't filter these out into a separate cost pool from the outset, you're trying to optimize a bill component you can't actually control.
The tagging effort should focus only on the remaining, truly tunable WAF rule triggers. Otherwise, you're building models on flawed data and will miss the real cost drivers.
CPU cycles matter
The timestamp formatting got us too. The API expects ISO 8601 but is picky about the timezone offset. We had to lock it down to Zulu time (ending with 'Z') to get consistent pagination.
On the silent field changes, that's a vendor problem. We started snapshotting the JSON schema from a known-good request as part of our CI checks. Any new deployment compares the live response keys against that snapshot and fails the build if fields are missing or renamed. It's blunt, but it catches those breaking changes before they corrupt a week's data.
For monitoring, we stuck with Python scripts feeding a simple time-series DB, but the key was adding that schema validation layer before the data ever hits storage.
Verifying "good" bots before the WAF is tricky. We run a small allowlist of IP ranges for known search engines and monitoring tools, but maintaining it manually is a pain.
The better approach for us was to use a lightweight, cheap CDN or proxy tier before the WAF. Incoming requests hit a super-basic rule there first: if the user-agent matches a known good bot pattern, we add a header and skip the main WAF entirely. It's not perfect, but it cut down on a lot of noise.
Did you find any automated signals, like verified reverse DNS, that were reliable enough to automate this?
pipeline all the things
That initial sticker shock is the alarm bell every team should listen to, but your "deep dive" into the billed volume is the only part that matters. You're on the right track dissecting the mitigated traffic, but the critical leap is realizing that a usage-based pricing model isn't a passive tax - it's an active incentive for the vendor.
Your "set-and-forget" security posture is exactly what they're banking on. When traffic grows, so does their bill, and their built-in defaults are rarely optimized for your wallet. The moment you start tagging rules and pulling reports, you're essentially doing cost-optimization work that should have been part of the initial configuration. The real question your team needs to ask is why legitimate bot traffic was ever hitting security rules in a way that incurred cost. Shouldn't the default posture for known-good actors be to bypass the expensive processing layer entirely?
Also, integrating the report into your existing monitoring is good, but I'd be wary of just feeding it into the same alerting pipeline. The metrics you're pulling are now directly tied to financial risk, not just system health. You need a separate model that treats cost-per-request as a service-level objective, because letting it drift is how you get another doubling next year.
Your k8s cluster is 40% idle.
The tagging and review process you implemented is essential, but its effectiveness depends heavily on how you define the baseline. You mentioned a "significant portion" was legitimate bot traffic hitting your security rules. Did your team quantify what percentage of the billed mitigated volume that represented versus network-layer DDoS actions? That split is the first calculation any cost analysis needs.
Your weekly script is the right tool. However, if it's not categorizing mitigations from the initial data pull - separating TCP resets from actual WAF rule triggers - you're still working with a blended, misleading metric. Rule tuning only impacts the latter portion.
What was the time delta between observing the bill increase and having the first actionable data from your script? In our case, that lag was the biggest budget killer.
Measure twice, buy once.
You hit on the crucial lag point. We saw the bill jump at the end of month one, but it took us three weeks to get the script stable and parsing correctly. That's an entire billing cycle of paying for noise before we could even start tuning.
Our split was eye-opening: about 40% of the "mitigated" volume was those non-tunable TCP resets. But your question about the baseline drift is spot on - we assumed that 40% was a constant ratio, but it wasn't. During a traffic surge the following month, the DDoS portion spiked to almost 60%, completely skewing our initial "actionable" calculations. The lesson was we had to make that categorization a real-time metric, not a one-time audit.
So the delta between seeing the cost and having clean data wasn't just a one-time lag; it created a moving target. If your script isn't continuously validating that split, you're optimizing yesterday's problem.