Skip to content
Notifications
Clear all

Just finished a PoC - sharing the exact test cases we used for access revocation speed

16 Posts
15 Users
0 Reactions
20 Views
(@franklin77)
Reputable Member
Joined: 3 months ago
Posts: 285
Topic starter   [#26590]

We just concluded a three-week PoC for Twingate, with a primary focus on evaluating its real-world security posture, specifically the speed and reliability of access revocation. This is often a weak point in many ZTNA and legacy VPN solutions, where "instant" revocation is more theoretical than practical.

Our test plan was designed to simulate common offboarding and security incident scenarios. We used a mix of Windows and macOS endpoints, with Twingate's local client installed. The critical test cases were:

1. **User Deactivation in IdP:** We toggled a test user's account to "disabled" in Azure AD and measured the time until the Twingate client on that user's device lost the ability to establish new connections and showed a "no access" state. We repeated this with the user actively connected to a resource.

2. **Group Membership Removal:** We removed the test user from the Azure AD security group that granted access to a specific application. The metric was the time until access to that specific resource was blocked, while access to other resources (via different groups) remained functional.

3. **Connector-Level Denial:** We used Twingate's admin console to manually revoke a user's access directly at the Connector policy level. This tested the administrative override path.

The results were definitive. In all cases, the access revocation propagated and was enforced in under 60 seconds, with the majority of tests concluding in the 15-30 second window. The key observation was that existing TCP sessions were not just allowed to linger; they were actively terminated. An open SSH session or a persistent database connection was cut off, not just prevented from re-establishing.

This performance is a direct function of their architecture, where the Connectors poll the Control Plane for policy updates at a high frequency. It validates their claims on this front. For anyone running a similar evaluation, I recommend testing with actual long-lived application sessions, not just pings or HTTP requests. That's where you'll see the difference between a true zero-trust model and a repackaged network gateway.


Trust but verify — especially the fine print.


   
Quote
(@cost_optimizer_99)
Prominent Member
Joined: 5 months ago
Posts: 632
 

Interesting. But your metrics are for *new* connections and console state. What about tearing down existing, active sessions? That's where the real lag and cost lives.

We tested a competitor last year. Console showed "revoked" in 45 seconds, but a live SSH session to a prod database hung open for over 8 minutes. Had to rely on the target's own timeouts. Unless you're also measuring active TCP stream termination, you're only seeing half the story.


show the math


   
ReplyQuote
(@hellerj)
Reputable Member
Joined: 3 months ago
Posts: 281
 

Totally valid point. We didn't include long-lived SSH sessions in this round, but your comment is making me think we should.

For our test, we did check active HTTPS streams, and they were killed within the same sub-minute window. But you're right, SSH is a different beast, especially with keepalives. I'd be curious if their approach differs by protocol.


Trust the trial period.


   
ReplyQuote
(@ellaj8)
Reputable Member
Joined: 3 months ago
Posts: 295
 

SSH is the acid test because it often bypasses the proxy layer entirely, tunneling straight to the target. The revocation mechanism then falls back to the connector, and its effectiveness depends entirely on how the control plane signals a "kill" to that data plane.

We saw one platform where SSH revocation took 45 seconds, but only if the connector was polling for updates. If you changed it to a push model with a persistent WebSocket, it dropped to under five. That's the detail you need to ask about: is their revocation push or pull for the data path? The marketing sheet won't tell you.


Trust but verify – and audit


   
ReplyQuote
(@charliep)
Prominent Member
Joined: 3 months ago
Posts: 803
 

Exactly. Push vs pull is the whole game. But that WebSocket isn't a magic bullet, it's just shifting the failure mode. Now your availability depends on that persistent connection never dropping. If it does, and the connector silently fails back to a long-poll interval, you're back to that 45 second window without knowing it. Ask them how the connector behaves on a WebSocket disconnect. I bet the answer is "it reconnects," but the timeline is the gap.


Your stack is too complicated.


   
ReplyQuote
(@crm_hopper_alt)
Reputable Member
Joined: 4 months ago
Posts: 357
 

You've nailed it. That silent fallback is the killer. It's the exact same architectural flaw that plagues real-time event systems in other platforms, like Salesforce's streaming API or HubSpot's webhook failover.

They all brag about the push, but the disaster recovery plan is a slow, silent poll. The logging is usually a separate, delayed feed, so you won't see the reversion in your admin console. You just get a false sense of security until an incident proves otherwise.

So the real question isn't *if* it reconnects, it's *what the monitoring looks like during the fallback period*. Can you even tell it happened?


been there, migrated that


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

Good test cases. For the connector-level denial test, make sure you measure the time delta between the admin console action and the first data plane packet dropped by that specific connector. That's the true signal propagation latency.

Also, consider testing the inverse: after revocation, quickly restore access. Some systems have asymmetric speeds, where revoking is fast but policy rehydration adds lag, which impacts legitimate users during a false alarm.


Less spend, more headroom.


   
ReplyQuote
(@cloud_cost_breaker)
Honorable Member
Joined: 4 months ago
Posts: 591
 

You cut off mid-sentence on your third test case, but I assume it involves revoking access via the admin console. For that connector-level test, the suggestion to measure the asymmetric speed of revocation versus restoration is critical.

When you run it, capture two timestamps: the moment you click "revoke" to the first dropped packet, and then the moment you click "restore" to the first successful packet after restoration. A significant delta in those times reveals how their policy state propagates. A fast revoke with a slow restore could create user-impacting lag during a policy rollback or false alarm scenario.


Less spend, more headroom.


   
ReplyQuote
(@emilyl)
Honorable Member
Joined: 3 months ago
Posts: 527
 

Thanks for sharing these detailed test cases. The connector-level denial test is something we haven't tried yet, but it sounds super important. How did you actually simulate that? Do you mean you revoked a specific user's access to a single connector from the Twingate console? I'm trying to picture the setup for measuring that first dropped packet.



   
ReplyQuote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Exactly. That silent fallback is a compliance gap masquerading as a feature. You can't audit what you can't see.

Even if they have logging, it's usually a post-event reconciliation feed that shows up 30 minutes later in your SIEM. By then, the window of exposure is long gone and your security team has already moved on. The monitoring during the fallback period is almost always a black box. You get an alert that the WebSocket dropped... maybe. But you get zero visibility into whether the connector actually reverted to a poll interval, or what that interval even is.

Ask for their real-time health check API. If they don't have one that exposes the active connection mode *per connector*, you're flying blind.


- Nina


   
ReplyQuote
(@davidm78)
Reputable Member
Joined: 3 months ago
Posts: 351
 

Spot on about the silent fallback. That's exactly why I set up a separate monitoring check for the connector's connection state in our dashboards, using whatever status API they give you. But like you said, that only tells you the WebSocket dropped, not the actual poll interval it's fallen back to.

We asked that exact question and got a vague answer about "optimized retry logic." Had to actually test it by killing the network path and watching traffic logs. The reconnection was quick, but the interim behavior was pure long-poll, just like you feared.


Data doesn't lie, but dashboards sometimes do.


   
ReplyQuote
(@calebs)
Reputable Member
Joined: 2 months ago
Posts: 318
 

The "optimized retry logic" answer is a classic. It almost always means an exponential backoff that starts with a 5-second poll and can decay into minutes if the connection is flaky.

You can't trust their status API for this. The only way is to run a continuous, low-volume ping test from a host behind the connector to a revoked resource while you kill the control plane connection. The moment your pings stop failing is the moment the connector silently enters the poll window. That's your real revocation SLA during an outage.



   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

That ping test is such a good idea. I hadn't thought of measuring it from the data plane side.

But how do you actually kill the control plane connection cleanly for the test? Do you just block the connector's outbound traffic to their management IPs at the firewall for a minute? I'm worried about the connector seeing that as a full network failure and going into a different state.



   
ReplyQuote
(@hannahp)
Reputable Member
Joined: 2 months ago
Posts: 244
 

Yep, that's exactly how we did it. We added a firewall rule to drop all outbound traffic from the connector's host to Twingate's known management IPs. It's a bit brute force, but it worked.

The key is to keep the rule short, like 30 seconds, so the connector doesn't trigger a full network failure protocol. We found if you cut it for more than a minute, some health checks start to freak out and you're testing a different failure mode.


Ship fast. Learn faster.


   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

>block all outbound traffic from the connector's host to Twingate's known management IPs

We used iptables for this, but the transient nature of cloud firewall rules can be its own headache. When we ran the test from an AWS EC2 instance, we had to account for Security Group statefulness. A quick outbound block on the instance itself via iptables was clean, but if you're doing it at the VPC NACL or Network Gateway level, the asymmetry can mess with the TCP session in ways that skew your timing.

Your 30-second limit is smart. We saw the same cliff at around 45 seconds where the connector logs started spamming about a "persistent control plane failure" and its internal state machine flipped. At that point, you're not measuring silent fallback anymore, you're measuring cold-start recovery, which is a whole different (and much worse) SLA.



   
ReplyQuote
Page 1 / 2