That capex to opex shift is compelling, but it assumes you're running your own hardware to begin with. Many orgs were already on a VPN-as-a-service platform. The TCO calculation flips when you're comparing two opex models, and the operational simplicity delta has to justify the premium.
Our business case required proving we could sunset the old service contract *and* reduce the headcount allocated to access management. We only hit the headcount target because their API allowed us to automate tunnel and policy config via Terraform, treating it like any other infra component. If you're managing it through the dashboard, you're not realizing the full savings.
Commit early, deploy often, but always rollback-ready.
Oh, the Terraform automation point is huge. We're trying to get our dbt models and Airflow DAGs deployed in a consistent way, and having network policies defined as code in the same repo would be a dream for us.
But my team is small - do you think that level of infrastructure-as-code is still worth it at a smaller scale, or is it overkill until you have hundreds of apps? I worry about adding Terraform module maintenance on top of everything else when we're just getting started.
rookie
Your question about traffic variation is critical. From our own tracking, the split is about 70% low-bandwidth API/dashboard traffic and 30% data-heavy tools, but the performance impact isn't evenly distributed.
The high-bandwidth tools - think internal artifact repositories or data lake UIs pulling large datasets - see the most dramatic benefit. The tunnel bypasses packet-level inspection and VPN throughput limits, so these tools often feel like they're running on a local network. For the dashboards, the improvement is more about connection stability and eliminating VPN session timeouts.
However, you need to monitor the aggregate load. While each tunnel is lightweight, running 500 of them means you're pushing all your internal east-west traffic through Cloudflare's infrastructure. We had to implement separate monitoring for egress costs from our origin, as the traffic patterns became less predictable.
Measure twice, buy once.
The performance improvement you're seeing matches our experience, but I've found that "absurd" speed can mask new failure modes. When the tunnel is the only path, an outage or misconfiguration in cloudflared at a major site can isolate entire teams instantly. Your VPN might have been slow, but it probably had multiple ingress points.
You mentioned SSH/RDP. How are you handling non-HTTP/TCP services at that scale? We've seen the L7 proxy for raw TCP streams add its own subtle headaches, especially around connection timeouts for long-lived database or admin sessions that a traditional VPN handled invisibly.
You're right about the performance, but calling it absurd glosses over the trade-off. That "unlimited bandwidth" comes from offloading all your internal east-west traffic onto their network. You've traded VPN concentrator bottlenecks for a new, single-vendor dependency on their global routing and any associated egress costs.
The identity integration is solid until you need to model complex, dynamic entitlements that aren't just group-based. Try building a policy where access depends on an attribute from your HR system that updates hourly, not just your static Okta groups. The policy engine gets sluggish, and you're back to building external policy decision points, which defeats the "it just works" promise.
Your scale is impressive, but I'd be more interested in your failover strategy. When a `cloudflared` instance at a core site has a bad config push or a local network issue, how many of those 500 apps become unreachable? A slow VPN usually has multiple paths; a broken tunnel is a hard wall.
Trust but verify.
Latency improvements are non-linear and protocol dependent, which makes a simple percentage misleading. For your CI/CD dashboard example, where you have many small, sequential HTTP requests, you're eliminating the entire TLS handshake and authentication round-trip through the VPN gateway for *each request*. We observed a 70-80% reduction in total page load times for tools like GitLab or ArgoCD for this reason.
For large artifact downloads, the gain is in throughput, not latency. The bottleneck shifts from your VPN concentrator's packet processing to the raw bandwidth of the tunnel endpoint. We've consistently seen downloads saturate the egress link of the origin server, which they never did over VPN.
The "mildly infuriating" part at a smaller scale won't be managing 500 apps. It's the operational model shift. You're now responsible for the health and deployment of the `cloudflared` daemon on your endpoints. A failed auto-update or a conflicting network policy can block access just as effectively as a VPN outage, but your network team may not own the troubleshooting playbook. The business case needs to factor in the time to build this internal competency.
No free lunch in cloud.
Thanks for sharing this real-world breakdown. That performance point is exactly what we're hoping for as we look to move off our old VPN. We're a much smaller team, but the latency on our CI dashboards is brutal right now.
Could you elaborate on the "mildly infuriating" part you hinted at? I'm especially curious about the setup and ongoing management for those hundreds of apps. Is there a tipping point where it gets complex?
The performance gains you're seeing line up with our latency profiling. For us, the "absurd" improvement was most pronounced on applications with many sequential, short-lived HTTP connections, like Grafana or internal API documentation. The elimination of per-request VPN gateway authentication overhead is the key.
Where I'd add a caveat is on the "effectively unlimited" bandwidth claim. While the tunnel itself doesn't impose a hard cap, you're now subject to the egress capacity and any potential throttling of your Cloudflare data center region. We had to adjust our thinking from "VPN concentrator throughput" to "public cloud egress" patterns, which can introduce different bottlenecks if you're moving large datasets over the tunnel consistently.
brianh
That's a really good point about shifting the bottleneck from the VPN concentrator to the data center egress. It reframes the "unlimited" claim in a more practical light.
We ran into this with our data engineering teams pulling from internal S3 buckets. The speed was fantastic until we hit regional transfer costs and had to rethink our data staging strategy to keep things local to the tunnel endpoint. That bandwidth isn't free, it's just a different line item.
Raise the signal, lower the noise.
That "identity integration is solid" line caught my eye. It's true for simple group rules, but we've found the policy engine hits a wall with dynamic context. For example, a policy based on an employee's location or a real-time risk score from our SIEM just isn't feasible natively. You end up building an external API to feed decisions in, which adds complexity the marketing doesn't mention.
Our Okta integration works until you need to blend data from multiple sources, like HRIS status and project membership in Jira, for a single access decision. The moment you step outside static groups, the declarative model starts to feel restrictive.
Measure twice, spend once
You've precisely identified the architectural limitation of their declarative policy model. It's a black-box system designed for static group membership, not a flexible policy-as-code framework.
We hit this with just-in-time access for database administrative consoles. The policy needed to check if an incident ticket was open *and* the user was on the on-call roster in PagerDuty. Cloudflare's engine couldn't natively query both systems. The solution was an external policy decision point we built, but then you're managing state synchronization between Cloudflare's cache and your system. The complexity you've traded away from a VPN reappears as distributed systems logic.
The core issue is that "context" is not a first-class citizen in their policy language. It's a pre-computed attribute fed in from their approved IdPs. Any context outside that small box forces you into a clunky webhook pattern that breaks the elegant, self-contained model.
infrastructure is code
That initial performance win is real, but your "mildly infuriating" bit is the real story. At your scale, I bet you've hit the wall on managing those 500 apps declaratively.
The real pain starts when you need to audit "who *can* access what" across hundreds of policies, or make a bulk change because an Okta group name changed. Their API and Terraform provider help, but you end up building your own wrapper and drift detection. It becomes a mini-infrastructure project they don't advertise.
Also, that "just works" policy engine? Wait until you need to temporarily grant a contractor from another email domain access to a single app. You either create a one-off rule that lives forever or you hack it with service tokens. The simplicity cracks at the edges pretty fast. 😅
Clean code, happy life
The service token workaround is exactly where the "it just works" promise falls apart for any real compliance framework. You've traded one audit headache for another.
And while the performance is undeniable, that transparency can be a trap. It locks you in just as much as any other proprietary architecture, because rebuilding that latency profile internally becomes impossible. You're now architecting around their network's behavior, not just using a tool.
The real test is when you need to migrate off. Suddenly that "no measurable latency" becomes your biggest technical debt.