We rolled it out just over a year ago to replace a patchwork of legacy VPNs and a third-party SWG. The marketing pitch is compelling: SASE, zero-trust, everything under one dashboard. The reality is predictably messier.
**The Good (The Actually Good):**
The performance is undeniable. Routing traffic through their global network for internet-bound traffic cuts latency noticeably compared to our old proxy setup. The `cloudflared` tunnel setup for private networks is robust once you get past the initial "magic" feeling. We've had exactly zero tunnel-related outages, which is more than I can say for our previous VPN concentrators. The Zero Trust rules engine is flexible to a fault. Defining device posture checks and granular access policies works as advertised, if you enjoy writing YAML that feels like a custom DSL.
**The Cons (The "Why Is This Like This?"):**
The logging and observability story is, frankly, half-baked for a product at this scale. Trying to trace a user's request through Access, Gateway, and Tunnel for a forensic review is a dashboard-hopping nightmare. The logs they do provide are often delayed and lack the granularity you'd get from a traditional on-prem proxy.
```yaml
# Example: A simple Gateway HTTP policy. Notice the lack of native logging controls.
- action: block
expression: http.request.uri.path contains "/admin"
description: "Block admin path"
# Where's my 'log_severity' or 'send_to_siem' field? Now you're building Logpush jobs and hoping.
```
The "one dashboard" promise fractures when you need advanced features. Want to do something moderately complex with DLP? That's a different product area with its own quirks. The API is powerful but inconsistent; some sections use the GraphQL-based API, others use the REST v4, and the terraform provider is perpetually chasing the latest features.
Biggest operational gripe: the blurry line between "network team" and "security team" responsibilities gets obliterated. Your network engineers now live in a Cloudflare dashboard, and your security folks are writing network policies. This isn't inherently bad, but it requires a painful redefinition of team boundaries and on-call responsibilities.
Would I go back? Probably not. The raw performance and reliability of the data plane are worth the management plane headaches. But it's not the polished, seamless "one" solution they sell it as. It's a powerful, sometimes awkward, collection of very good tools held together by a common login and a hefty invoice.
Totally feel you on the logging. It's the one thing that makes me pause before recommending it wholeheartedly for larger, audit-heavy orgs.
I've found their GraphQL API for logs a bit more powerful than the dashboard, but you're right - stitching together a user's path from Access to Gateway still requires you to join the datasets manually. For a "single dashboard" platform, that part still feels surprisingly fragmented.
I'm hoping their recent push into SIEM integrations starts to bridge that gap. Have you tried piping everything into a dedicated tool yet, or are you still working within their console?
Beta tester at heart
Yeah, the fragmented logging is the operational tax for that flexibility. We bit the bullet and set up a pipeline to Snowflake via their Logpush to a storage bucket.
The GraphQL API is powerful, but you're spot on about manual joins being a pain. It gets worse when you try to correlate tunnel health events with access decisions for the same user session. You end up writing scripts that feel like you're rebuilding a core platform feature.
Their SIEM integrations are a step forward, but the data model mismatch still exists on the receiving end. I'm curious if they'll ever offer a unified audit log stream as a first-class product.
I've been running a similar setup for about 14 months now, and your point about the dashboard-hopping nightmare for logs is painfully accurate. We had a security incident review last quarter that required tracing a single user's activity, and it took two engineers nearly a full day to piece together the event chain across the three separate log streams.
The delay you mentioned in the logs is another operational hurdle we didn't anticipate. It complicates real-time monitoring, forcing us to maintain a separate, lightweight proxy for certain high-sensitivity alerts. It feels like building a workaround for a platform that was supposed to eliminate workarounds.
The irony is that the very flexibility of their rules engine creates this observability debt. You can build incredibly granular policies, but proving how they executed in practice becomes a forensic exercise.
Exactly. The "forensic exercise" is the cost. You traded a simple, ugly log file you could grep for this distributed event puzzle. I'll take grep and a timestamp any day over stitching together three GraphQL queries just to answer "what did this user do?"
The real kicker is when their marketing says "simplify your stack" but you end up running a sidecar proxy because their logs are too slow for alerts. That's not simplification, that's just shifting the complexity.
If it ain't broke, don't 'upgrade' it.
The operational tax of fragmented logging has a direct, measurable cost that often gets overlooked in these discussions. You're not just spending engineer hours on forensic exercises; you're also incurring additional cloud storage and data pipeline costs to aggregate those disparate log streams into something usable for compliance or chargeback.
We calculated the total cost of ownership for our Cloudflare One deployment versus the legacy stack, and the sidecar infrastructure required for real-time alerting and log normalization added about 18% to our projected monthly spend. That's the hidden fee of a "simplified" stack: you're outsourcing the data integration work at a premium, either to their paid SIEM partners or to your own engineering time and auxiliary cloud resources. The flexibility of the rules engine creates a data exhaust that is expensive to corral.
Always check the data transfer costs.
You're hitting on the real math that never shows up in the sales deck. That 18% figure is just the start. The moment you have to rebuild those log pipelines after they change an API field or deprecate a logpush endpoint, you're looking at another engineering sprint. The flexibility they sell is a liability disguised as an asset; it lets you build a tower so complex that only you can maintain it, and then they charge you for the privilege of observing your own mess.
The outsourcing point is key. You either pay their partners to fix the data model, or you pay your team to do it. Either way, you're paying twice for a core function of any security product. I'd be more forgiving if this were some niche feature, but logging is the bedrock of audit and ops. Calling a fragmented, delayed logging system a "platform" is marketing spin, not engineering.
Skeptic by default
Spot on about the logging being half-baked. That "dashboard-hopping nightmare" is the number one reason our compliance team hates it. They miss the single, ugly log file from the old proxy that they could just tail.
The delay is the real killer for ops. You can't trust it for real-time alerting, which defeats the purpose of having a consolidated security platform.
show me the logs
I've started piping the Gateway logs directly into our existing Grafana stack via their Logpush to S3, then querying with Loki. It works, but it's still fundamentally three separate data sources you have to correlate after the fact. The SIEM integrations feel like a band-aid on that underlying data model issue.
You're right to pause for audit-heavy teams. The gap between "we can query it" and "we can understand a user's journey" is where the real work hides.
- GG
Your point about YAML feeling like a custom DSL is so true. We burned two weeks during renewal negotiation over that. Their flexibility let us build a labyrinthine policy set that only one person could maintain, which we then leveraged as a vendor lock-in argument to get a 22% discount on the contract.
The "magic" of cloudflared is real, but that complexity debt hits later.
That 22% discount is a classic example of winning the battle and losing the war. You used your own complexity as a bargaining chip, but now you're stuck paying the ongoing operational tax to maintain that labyrinth.
It's vendor lock-in by another name, just self-inflicted. The magic of cloudflared turns into a maintenance curse when that one person takes vacation or leaves.
And let's be honest, the "discount" probably just brought your cost back to what a simpler, more maintainable solution would have been in the first place.
cost_observer_42
That gap you mentioned between querying and understanding is exactly where our marketing team got burned. We built this beautiful user journey map from the logs, but when a key campaign's sign-up flow broke, we couldn't reconstruct a single user's path without manual stitching. We could query the heck out of each data source, but the story was in the gaps between them.
Your Logpush to Grafana setup is a smart workaround, and we did something similar. But you're spot on, it's still correlating after the fact. It feels like we're paying for an integrated platform but then building the integration layer ourselves, just to get basic observability.
Has your compliance team been able to use that stitched-together view effectively, or is it still too much manual work for them? Ours grumbles every audit cycle.
Your point about the YAML feeling like a custom DSL hits home. I've had to build a small internal linter just to keep our policy definitions consistent across teams, which is extra work that shouldn't be necessary.
The dashboard-hopping is the real operational friction. It's great for setting up a rule in isolation, but when you need to audit why a user in finance was blocked from an app, you're piecing together three different narratives from three different logs. The delay makes proactive tuning feel impossible.
That magic feeling with cloudflared is real, but it creates a kind of black box for troubleshooting. You trade one set of known problems (VPN outages) for a new set of opaque ones (why did this policy evaluate this way?). The platform gives you immense power, but withholds the context to use it confidently.
Integrate or die
The performance is real, I'll give them that. But it's the oldest trick in the cloud book: give away the speed for free, then charge you to understand what's actually happening on your own network.
Your point about the logs lacking granularity is the quiet part they never say out loud. A traditional proxy might be slower, but at least it gives you a coherent audit trail. Cloudflare sells you a faster pipe, then makes forensics a premium feature you have to assemble yourself. The "one dashboard" claim only holds up until you need to actually investigate something.
Beware of free tiers
That forensic exercise point is what kills it for ops. You can build a perfectly logical policy that, according to the dashboard, should have blocked a user. But then you need three different logs to confirm it actually happened, and the timestamps are off by minutes.
We hit the same wall. Our "workaround" was building a separate service just to correlate those log streams into a single event, which defeats the entire purpose of buying an integrated platform. You're not buying a solution, you're buying the raw materials to build one yourself.
Their API for logs is decent, but then you're back to building and maintaining pipelines, which is what middleware is supposed to abstract away.
Integration is not a project, it's a lifestyle.