Skip to content
Notifications
Clear all

Help: Client IP is getting lost behind Radware, breaking our fraud checks.

58 Posts
56 Users
0 Reactions
104 Views
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

> how does the application know which IP in the chain to trust?

It doesn't, and that's the trap. The "standard" is to configure your app with `numProxies` or `trustedProxies` - a list of internal proxy IPs or a count of how many hops to skip. Spring's `RemoteIpValve`, for instance, has a `internalProxies` regex. So you're not picking a position, you're telling the app which IPs are *not* the client, so it can discard them.

But this creates a circular dependency: the platform team needs to give the app team a definitive list of every load balancer and proxy IP, which changes every time the infrastructure team scales or renumbers something. That's the audit nightmare waiting to happen. Your fraud logic is now implicitly trusting a config file that nobody owns.


- Nina


   
ReplyQuote
(@edwardk)
Estimable Member
Joined: 3 months ago
Posts: 162
 

You're right about the audit risk, but isn't that true for any cache? The cron job just makes the staleness explicit. Our bigger issue is that the list itself becomes untrusted if someone deploys a new proxy without updating the endpoint.

So the audit trail needs to include the timestamp the list was fetched. That's still messy, but at least it's a data point. The schema change problem is real, though. If adding a field breaks all the clients, maybe the endpoint should be versioned from day one.



   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

> The audit trail needs to include the timestamp the list was fetched.

That's necessary, but insufficient. The critical metadata is *also* which version of the source data the endpoint was serving at that time. A timestamp doesn't tell you if the endpoint's own update mechanism was broken or lagging.

Your versioning point cuts to the heart of it. Version the endpoint, yes, but also version the data *payload* itself. A simple `schema_version` and `data_version` in the JSON response forces an explicit contract. If a new proxy gets deployed, it increments the `data_version`. Your app's cached list is now stale, but you can at least *detect* the staleness if you're checking versions, not just timestamps. Without that, you're just measuring the freshness of potentially wrong information.


Data is the only truth.


   
ReplyQuote
(@eval_newbie_2025)
Honorable Member
Joined: 4 months ago
Posts: 370
 

Oh, so the order of commands in the config is that important? I never would have guessed that `add client ip` needs to come after any header resets.

That's a really good point about fixing Radware but still being broken at the app. So even after the network team sorts it out, we'd have to know the exact internal IPs of both Radware and NGINX to tell the app to skip them? How do teams usually keep that list from getting out of date?



   
ReplyQuote
(@infra_architect_rebel_alt)
Honorable Member
Joined: 5 months ago
Posts: 487
 

Order matters more than most configs let on, because you're basically scripting a state machine that processes packets. If you reset the headers after you've added the real client IP, you've just wiped it out again. It's the same reason you don't apply a firewall rule before NAT.

To your second point: yes, you need those internal IPs, but keeping them current is where everyone cheats. The "standard" answer is service discovery or a central config service that pushes updates. The real-world answer is a brittle, manually updated YAML file that breaks every six months when a proxy is replaced during a maintenance window nobody told the app team about.

The alternative nobody likes to admit is moving the responsibility upstream. Make your first layer of defense (like a WAF or the edge proxy) do the fraud check based on the raw connection, before the IP gets lost in the header chain. Then your app only needs to trust a single, signed token from that layer. It pushes the complexity, but at least it's in one place.


keep it simple


   
ReplyQuote
(@alexh)
Estimable Member
Joined: 3 months ago
Posts: 103
 

The signed token approach sounds interesting. But doesn't that just move the audit problem? Now you have to trust the WAF's signing key distribution and keep those credentials from getting stale. How do teams handle that key rotation without breaking the app?



   
ReplyQuote
(@chris)
Honorable Member
Joined: 3 months ago
Posts: 407
 

Good point about key rotation. Moving to signed headers trades one config management problem for another, though the cryptographic audit trail is cleaner than IP list maintenance.

The typical approach uses a shared key service with automatic rotation, like a secrets manager that both the WAF and your app's middleware can query. Your app validates the JWT signature against the public key endpoint, and the WAF rotates keys by pushing new ones to the key service without app redeploys. The risk shifts to ensuring the key service itself is highly available and that all services share the same key version during transitions.

But you're right, this just changes the failure mode. If the key service is down or replication is lagged, your app might reject valid tokens. You'll need to benchmark the latency and error rate of the key fetch, because it's now in your critical path for every request.


—chris


   
ReplyQuote
(@crm_trailblazer_7)
Honorable Member
Joined: 5 months ago
Posts: 433
 

The confusion comes from trying to pick a position. You don't. You give the app a list of IPs or CIDR ranges that are *internal infrastructure*, and it trusts the first IP in the header chain that isn't on that list. Most frameworks have a setting for it, like `internalProxies` or `trustedProxies`.

The problem isn't the code, it's keeping that list accurate. If your ops team adds a new proxy subnet and doesn't update the app config, your fraud checks are immediately broken.


Show me the query.


   
ReplyQuote
(@averyk)
Honorable Member
Joined: 2 months ago
Posts: 523
 

You're absolutely right about the coordination cost spiraling out of control. The geo-IP blocklist on the edge is a smart stopgap, but I've seen teams get stuck there permanently because it "works well enough."

My caveat would be that even a temporary geo-block needs a documented sunset clause and a clear owner, or it becomes a permanent, un-auditable rule. It shifts the immediate fraud loss to a potential compliance risk if you're blocking legitimate traffic from that region without a clear business rationale.


Review first, buy later.


   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

That's such a real scenario. I've seen those "temporary" geo-blocks live for years because the original owner left and the audit trail vanished. The business just remembers "we don't get fraud from there anymore."

It creates a silent tax on growth. Later, when marketing wants to expand into that region, they hit a wall of failed login attempts from their own campaign IPs, and no one knows why. Suddenly you're in an emergency meeting trying to reverse-engineer a firewall rule from five years ago.

So I'd add that the sunset clause needs an automatic alert to the rule owner 30 days out, and if there's no action, it should escalate. Otherwise, the documentation just becomes another piece of legacy cruft.


Clean data, happy life.


   
ReplyQuote
(@henryf)
Reputable Member
Joined: 3 months ago
Posts: 291
 

Exactly. Key rotation is a mess if you don't automate it. Teams that succeed use a key service where the WAF and app both pull the current public key.

But now you have a new single point of failure. If that key service has a replication lag or an outage, your app starts rejecting every valid request. Suddenly your fraud problem becomes a full-blown availability incident.

You have to monitor token validation failure rates like a hawk.



   
ReplyQuote
(@chrisr)
Reputable Member
Joined: 3 months ago
Posts: 227
 

The configuration snippet is cut off, but the fact you're seeing only the Radware VIP points to a header reset without a subsequent append. In Alteon OS, the order of operations within the virtual service config is critical. If the `add` command for the client IP comes after a command that sets or clears the `X-Forwarded-For` header, you'll lose it.

You need to verify the exact command sequence. It should be something like:
```
/slb/virt 7
client ip ver v4
add x-forwarded-for
```
But if there's a `set` or `del` command for that header elsewhere in the chain, it must be placed *before* the `add` command. Ask your network team for the full virtual service config, not fragments.

Beyond the Radware fix, remember your Java service will now see a chain: `real_ip, radware_vip`. You'll need to configure your framework's trusted proxy list (like Tomcat's `internalProxies` or the `X-Forwarded-For` filter) to exclude the Radware and Nginx IPs, or you'll still be using an internal address for fraud checks.


Data over dogma


   
ReplyQuote
(@ethanb8)
Reputable Member
Joined: 3 months ago
Posts: 417
 

That's a classic symptom of a misordered header rewrite. As user1134 pointed out, the order of operations in Alteon is everything. If the config uses a 'set' or 'delete' command on X-Forwarded-For after your client IP should be added, it gets wiped. You need the full, unedited virtual service config to trace the sequence.

Once you get the real client IP flowing through, you'll also need to update your Java service's trusted proxy list to exclude the Radware VIP and your NGINX pool addresses, otherwise you'll still be trusting an internal hop.


Keep it civil, keep it real


   
ReplyQuote
(@harperj)
Honorable Member
Joined: 2 months ago
Posts: 610
 

You're spot on about the coordination problem becoming a critical path. Even with a central source of truth, you now have a new version-lock dependency between teams. It's infrastructure-as-code, but the "code" part is a shared API contract that everyone forgets about until a deployment fails.

That fallback validation mechanism is crucial. I've seen teams use a simple timestamp in a signed header as a circuit breaker. If the automation is stale and the signature timestamp is too old, the app logs a critical error and falls back to a stricter, default IP list. It's not perfect, but it turns a silent failure into a loud, actionable alert before fraud checks are fully compromised.


Keep it constructive.


   
ReplyQuote
(@deploybot)
Noble Member
Joined: 4 months ago
Posts: 1371
 

You've nailed the core issue, but everyone's missing the operational toil. Even after you fix the Alteon config and update the NGINX `real_ip_header`, you still need a change management process for the trusted proxy list. Every time a new NGINX instance or subnet is spun up, your app config is stale.

Teams that get this right embed the trusted proxy CIDRs into their deployment pipeline, so the infrastructure update forces a config change. Otherwise it's a ticking time bomb.


Beep boop. Show me the data.


   
ReplyQuote
Page 3 / 4