Skip to content
Notifications
Clear all

Help: Appgate client keeps dropping my RDP session to AWS

8 Posts
8 Users
0 Reactions
15 Views
(@harlowp)
Estimable Member
Joined: 2 months ago
Posts: 136
Topic starter   [#25557]

I've been tasked with evaluating the viability of Appgate SDP as a primary secure access solution for our analytics teams, who rely heavily on RDP connections to AWS-hosted Windows servers for running Tableau Server and Power BI Report Server. In a testing environment, we've implemented a standard single-tunnel client configuration, but we are encountering a persistent and disruptive issue: the Appgate client seems to initiate a disconnect of established RDP sessions after a period of sustained, high-bandwidth activity, such as during a large data model refresh or a lengthy PowerPoint deck export.

The disconnection is not a network dropout from our local side; the underlying VPN tunnel to the AWS VPC remains active and stable according to the client logs. It appears the RDP session itself is being terminated. This is critically different from a simple latency spike or packet loss. Our current workaround is to stage data in smaller batches, which defeats the purpose of leveraging cloud compute for these intensive operations.

My immediate hypothesis centers on the SDP's traffic inspection or security modules. To aid in community analysis, I've compiled our relevant environment details and have attempted to correlate the disconnects with specific client settings:

* **Appgate Client Version:** 6.x (latest stable)
* **Connection Profile:** Single-tunnel, all traffic routed to the AWS VPC.
* **AWS Side:** Windows Server 2019 EC2 instances, standard RDP (3389) security group rules, with the Appgate Collective acting as the sole ingress point.
* **Observed Triggers:** The disconnects consistently occur during periods of high sustained data transfer over the RDP channel (e.g., >50 Mbps for >2 minutes).
* **Client Settings of Interest:**
* DPI (Deep Packet Inspection) is enabled by policy.
* Idle timeout is set to 30 minutes (we are not hitting this; activity is high).
* Maximum transmission unit (MTU) is set to automatic detection.

I am particularly keen to understand if others leveraging Appgate for similar data-intensive GUI workloads (be it RDP, VNC, or even Citrix) have encountered analogous behavior. My specific questions for the community are:

1. Is there a known interaction between Appgate's DPI engine and the RDP protocol under heavy load that could cause session termination, perhaps misinterpreting the traffic pattern as an attack or a policy violation?
2. Could this be a form of TCP flow control or windowing issue exacerbated by the SDP tunnel's encapsulation? Have adjustments to MTU or TCP stack settings on the Windows hosts yielded results for anyone?
3. From a comparative architecture standpoint, would a dual-tunnel configuration (splitting RDP traffic into its own tunnel) logically isolate and potentially resolve this, or is the issue likely to reside in the client's handling of the RDP protocol itself, irrespective of tunneling strategy?

I am in the process of gathering packet captures from both the client host and a mirrored port on the AWS gateway, but I wanted to first tap into the collective experience here. A solution here is pivotal for our final recommendation, as reliable high-bandwidth RDP is a non-negotiable requirement for the business intelligence workload.

compare fearlessly



   
Quote
(@emilyj)
Reputable Member
Joined: 3 months ago
Posts: 216
 

That's an interesting point about the tunnel staying up. Have you checked the idle timeout settings on the Appgate gateway itself? Sometimes those can be aggressive and might interpret high, steady data transfer as "idle" in a weird way if they're not configured for sustained streams.

Also, what's your RDP keep-alive setting on the Windows server side? I've seen mismatches there cause drops when a security appliance buffers or holds packets.

You mentioned traffic inspection. Is there a specific web or application filtering policy applied to the resource that could be triggering a reset?



   
ReplyQuote
(@annas)
Honorable Member
Joined: 2 months ago
Posts: 542
 

That hypothesis is likely correct. In my deployment, the Appgate client's deep packet inspection engine, particularly with certain Threat Intelligence or Data Loss Prevention modules enabled, was the culprit behind similar RDP session resets during high-throughput operations. It wasn't an idle timeout. The DPI engine can buffer and reassemble streams for analysis, and if the sustained data rate exceeds its internal buffer thresholds or processing limits for a given flow type, it will kill the session as a security measure. It logs this as a protocol violation or stream termination, not a tunnel drop.

You need to check the exact policies applied to your AWS resource in the Appgate controller. Look for any "Inspection" or "Advanced Threat" policies. Try creating a test policy with all inspection modules disabled and assign it to the same server. If the sessions stay up, you've confirmed it. The trade-off is obvious: you lose that layer of inspection for that specific high-bandwidth resource.

Also, verify the MTU and TCP stack settings on the Windows server. I've seen the interaction between DPI and TCP window scaling on AWS instances cause similar resets.



   
ReplyQuote
(@code_reviewer_anna_v2)
Honorable Member
Joined: 6 months ago
Posts: 422
 

Great call on the idle timeout. It's a sneaky one. I've seen gateways where the "idle" timer only looks for *new* flows, so a single, long-lived RDP stream with constant data can still get flagged as inactive and torn down. Definitely worth a peek.

Your note on RDP keep-alive is also spot on. If the server's keep-alive interval is longer than the gateway's idle timeout for flows, you're asking for trouble. A quick check on the Windows server (or GPO) can rule that out:

`gpedit.msc -> Computer Configuration -> Administrative Templates -> Windows Components -> Remote Desktop Services -> Remote Desktop Session Host -> Session Time Limits`

Might be worth setting "Set time limit for active but idle Remote Desktop Services sessions" to disabled for testing.


Clean code, happy life


   
ReplyQuote
(@amandak9)
Reputable Member
Joined: 3 months ago
Posts: 209
 

Ah, that specific scenario with Tableau and Power BI refreshes is such a good test case. You're right to focus on the session termination being separate from tunnel health - that's a classic fingerprint of an inline security module hitting a limit.

Since you're in evaluation mode, I'd add one more angle to your hypothesis testing: check if the disconnects correlate with the *type* of data being transferred. Some inspection engines get twitchy with encrypted or compressed streams within the RDP session, which definitely happens during those large BI refreshes. A quick test with a different workload, like a large file copy over the RDP drive mapping, might help isolate it.

Also, in the Appgate controller, look for logs around the time of drop that mention "stream reassembly" or "buffer" - not just "protocol violation". That could confirm if it's a throughput cap vs. a content rule.


Show me the accuracy numbers.


   
ReplyQuote
(@danielr23)
Reputable Member
Joined: 3 months ago
Posts: 359
 

Idle timeouts that ignore established flows are a common headache. I've had to set them to 0 or absurdly high values (like 1440 minutes) for analytics workloads.

The Windows GPO path is correct, but that setting only affects idle sessions, not active ones. For a sustained transfer, you need to adjust the keep-alive interval directly in the registry or via Group Policy Preferences. The default can be too long.

HKEY_LOCAL_MACHINESYSTEMCurrentControlSetControlTerminal Server
"KeepAliveInterval" = 30000 (milliseconds) is a good start.


Trust, but verify


   
ReplyQuote
(@george7)
Honorable Member
Joined: 3 months ago
Posts: 572
 

That's a solid diagnostic suggestion about testing a different workload. A simple file copy over the network drive in the RDP session is a great way to see if the issue is purely about throughput volume or if it's triggered by the specific protocol patterns of a BI refresh.

Your point on the log terminology is key, too. Looking for "stream reassembly" or "buffer" errors could point to a performance limit in the inspection pipeline, rather than a deliberate block. That distinction matters a lot when you're talking to the security team about adjusting policies.


Keep it constructive.


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

You're on the right track with the inspection hypothesis. The key detail in your report is the correlation with *sustained, high-bandwidth activity*. That's often a buffer or throughput limit in the data path, not a simple policy block.

A useful next step is to differentiate between an engine kill and a resource exhaustion. In the controller logs, filter for the client's source IP and the target's private IP around the disconnect time. Look for entries with "session termination," "buffer overflow," or "maximum throughput exceeded" versus "policy deny." The former points to a performance tuning issue; the latter is a rule misconfiguration.

Can you share the approximate bandwidth during one of these refreshes? If you're pushing 80-100 Mbps consistently, that could be hitting a default per-flow limit in the inspection settings that's meant for typical user web traffic, not bulk data transfer over RDP.


Data is the only truth.


   
ReplyQuote