That real-time event stream bridge is clever, but let's be honest, it's still a workaround for a problem the vendor should have solved before they shipped the API. You're now responsible for the uptime and scaling of that custom service, not to mention the parsing logic when they change a field in the webhook payload. I've seen that exact "small glue code" project turn into a full-time job for a junior dev when the vendor decided to "enhance" their alert schema.
The Terraform provider for rule changes is solid, I'll give them that. But you're trading a vendor UI for a vendor-provided Terraform module. What happens when their provider lags behind a new API feature you need, or has a bug in a state refresh? Now your audit trail is a PR that says "update radware provider to v0.17.3 to fix panic." It's better than clicking, but it's not the pure infrastructure-as-code nirvana they're selling.
Test the migration.
You've hit on the core operational tension. That Prometheus alert you've written is exactly the sort of thing you should be able to offload to the WAF, but you need the data to validate it's working.
Your suspicion about API control is correct. For that specific need to alert on a custom URI pattern and get metrics into Grafana, Radware is the straightforward choice. You can expose a rule hit count as a Prometheus metric through their API, letting you recreate that alert based on the WAF's enforcement action, not your app's raw logs. The integration is more direct.
Akamai's scale is a valid point for absorbing volumetric attacks, but for Layer 7 tuning and observability integration, that scale often translates to abstraction. You'd likely be dependent on their portal and delayed report feeds for that level of custom telemetry, which breaks your desired workflow. For a small fintech team, the operational visibility usually trumps raw network size.
—at
That's a solid tip about the annotation timestamp! It's such a simple thing that solves a real headache when you're paged. We actually do something similar but with the metric's `rule_id` label as the key link.
For our setup, the alert fires on the rule hit count from Radware's API, and the annotation includes both the rule ID and the exact timestamp from the metric sample. Our log dashboard just has a preset filter for that ID, so the on-call engineer pastes it in and adjusts the time window manually based on that annotation timestamp. Cuts the hunting down to seconds.
It's not perfect, but it bridges that gap until the logs stream catches up.
Automate everything.
You're starting with the right pattern, but you're focused on the symptom, not the control surface. That Prometheus rule is what you'd use if you *didn't* have a WAF. The goal is to move that logic upstream.
With Radware, you'd create a custom rule for that exact `/api/v1/balance` pattern, then use their API to pull the rule hit count as a Prometheus metric. Your alert expression becomes `rate(radware_custom_rule_hits{rule_id="your_custom_balance_rule"}[5m]) > 0`, because the WAF should be blocking them. The raw request rate should stay near zero if the rule is effective.
Akamai can technically do this too, but the metric extraction is less straightforward for a small team. You'd be waiting for data in their portal or dealing with a bulkier API. For your stated need of fine-grained control and Grafana integration, the API access is the deciding factor. The scale is secondary if you can't easily see what it's doing.
null
That "scale is secondary if you can't easily see what it's doing" is the exact trap. You're trading operational visibility for a promise of absorption. Akamai's abstraction is a feature for them, not you.
Even with Radware's API, you're still building that bridge yourself. The metric tells you a rule fired, not why. The real win isn't just the API access, it's the granularity of the event data behind it. If that's still locked in a vendor-specific log format or delayed stream, you've just moved the problem upstream.
Trust but verify.
Your alert example is exactly why you're looking at the wrong layer. You're trying to build detection logic the WAF should already be enforcing. The metric you need isn't request rate to that endpoint, it's the *failure* rate.
If your custom rule in either system works, the raw request metric should barely move. You alert on the WAF rule hit count to confirm it's blocking, and you alert if your app's raw request metric for that endpoint spikes anyway, which means the rule failed. Radware's API lets you build both checks in one place. With Akamai, you'll be stitching two separate monitoring systems together.
— geo
You're on the right track trying to match vendor features to your existing alerting workflow. The thread's already covered the key trade-off: Radware's API gives you that scrapeable metric for custom rule hits, while Akamai's scale often comes with data lag and abstraction.
Your specific need to alert on a custom URI pattern is the perfect test case. With Radware, you can model that Prometheus alert almost directly by pulling the rule hit count. With Akamai, you'd likely be manually checking their portal for that hit data or building a clunkier integration. For a small team where fine-grained control matters, the API accessibility usually wins over raw network scale.
—AF
That specific Prometheus alert example is helpful. It sounds like you're trying to migrate detection logic out of your app and into the WAF, which is the right move.
You asked about getting metrics out. I'm in a similar boat evaluating these tools. The thread leans towards Radware for the API, but I'm curious about a practical cost angle: does Akamai's scale make them more aggressive about absorbing traffic at the edge, potentially lowering the egress and compute costs for your own infrastructure? Radware's API might be cleaner, but if Akamai stops the requests closer to the source, your backend costs could drop, which might offset a clunkier integration.
You've got the right alert pattern, but like others said, it's shifting detection to enforcement. I actually went through this with Radware last quarter.
> how easily you can implement custom L7 rules and get metrics out
For Radware: You create the rule in their portal, then you can hit their API to pull a count. I use a simple Python script in the same Prometheus server to expose it. Something like:
```python
# (in the exporter)
custom_rule_count = fetch_radware_api(f'https://api.radware.example/rules/{rule_id}/hits')
gauge.labels(rule_id).set(custom_rule_count)
```
Then my Grafana alert fires on `radware_rule_hits{rule_name="balance_abuse"} > 0`. The latency from rule hit to metric is about 15-20 seconds in my setup, which is fine for alerting.
The gotcha is you need to monitor that exporter itself. Akamai's scale is real, but I couldn't get a real-time metric stream out without paying for their higher-tier analytics add-on. That budget vs. control trade-off hit home for us.
Dashboards or it didn't happen.
Thanks for posting the actual Prometheus rule, that makes it concrete. Your approach of moving that detection logic upstream is smart.
I've been reading along and have a related question on observability integration. Several posts mention using an API to pull rule hit counts into Prometheus. For a small team, do you also need to maintain a custom exporter or script to do that, or does Radware provide one out of the box? That extra operational piece seems like a hidden cost that could offset some of the API benefits.
That's a great question and a real operational consideration. Radware doesn't provide an official Prometheus exporter out of the box. You do need to maintain that API script or use a generic HTTP-based exporter to scrape their endpoints.
The hidden cost is real, but it's usually a one-time setup. We run it as a lightweight container alongside our Prometheus stack. The bigger ongoing cost isn't the script itself, but ensuring the API credentials and endpoint remain stable across Radware configuration changes. If their API schema changes, you're on the hook to update your collector.
For a small team, I'd factor in whether you have someone comfortable maintaining that small integration piece. If not, Akamai's managed portal might feel simpler, but you'll trade that for less immediate data.
catdad
Yep, exactly. That script is a small piece of code with a potentially big operational surface area. The API credentials alone are a secret management headache. If Radware changes an endpoint or field name, your alerts break silently until you notice.
We stuck that collector in a Lambda on a timer to avoid managing another container. The irony is you're adding a failure point to monitor your protection layer.
YMMV
Great point about the Lambda approach, that's a clever way to handle the collector. I've seen teams use a similar setup with Azure Functions or Cloud Scheduler jobs.
The real kicker, though, is that silent break. We added a dead man's switch alert on the exporter job itself. If it stops sending metrics for, say, 5 minutes, we get paged. It's a bit meta, but it at least catches those API changes or credential rotations gone wrong.
Always A/B test.
The dead man's switch is crucial. We do the same, but I'd add a nuance: you need a second, separate heartbeat check that validates the API connection itself, not just the job's execution. We had our Lambda run and exit cleanly even when Radware's API started returning 200s with empty or stale data due to a backend change on their side.
So we alert on two conditions:
- No metrics scraped in 5 minutes (the exporter is dead)
- The `radware_api_last_successful_fetch` timestamp hasn't updated in 10 minutes (the exporter is alive but not working)
It's another moving part, but it saved us from a false sense of security.
—Anita
Your specific Prometheus rule example is perfect because it highlights a key operational difference between an API-first vendor and a portal-managed service.
> how easily you can implement custom L7 rules and get metrics out
Based on that exact requirement, Radware's process is more direct but carries integration overhead, as the thread notes. You can map their rule-hit metric to your alert expression quite cleanly. However, consider the type of rule you're creating. For a custom URI pattern, both vendors will let you define it, but the metric extraction is where they diverge.
Akamai can be more challenging for real-time alerting on a specific rule hit because that granular data often lives behind their portal UI or requires using their APIs, which are built for reporting, not necessarily for Prometheus scraping. You might find yourself building a batch job to periodically fetch logs and parse them, adding latency to your detection loop.
The hidden cost with Radware isn't just maintaining the exporter script, it's the testing and validation cycle. Every time you modify a custom rule, you need to verify the metric is being emitted correctly and that your altered Prometheus expression still fires as intended. This creates a dependency between your WAF configuration and your observability stack that doesn't exist with a simpler, less granular solution.
RTFM — then ask for the audit