Skip to content
Switched from using...
 
Notifications
Clear all

Switched from using Claw's cloud to their on-prem version. The operational burden is insane.

3 Posts
3 Users
0 Reactions
19 Views
(@bearclaw)
Reputable Member
Joined: 3 months ago
Posts: 397
Topic starter   [#27904]

They said it was a "self-hosted" solution. What they meant was we now host their entire engineering team's job security.

The promised "single-binary deployment" turned into a 47-step Ansible playbook that assumes a pristine, air-gapped Kubernetes cluster from 2021. Their "lightweight agent" requires a custom kernel module. Don't even get me started on the "highly available" control plane—it's a three-node etcd cluster that throws a fit if clock skew exceeds 50ms.

```yaml
# Their example config for "basic" monitoring.
claw_onprem:
dependencies:
- cassandra: ">=3.11, <4.0"
- rabbitmq: "3.8.12 exactly"
- legacy_zookeeper: "3.5.9"
resource_requirements:
minimum_viable_cluster: 32 cores, 128GiB RAM
recommended_for_production: your entire data center
```

The cloud version was a dream. This is the part where you wake up and find you're now responsible for patching, scaling, and debugging their entire stack. The bill might be lower, but they've just outsourced their ops to you. At your own day rate, you're losing money by hour two.


Prove it.


   
Quote
(@elenar)
Reputable Member
Joined: 3 months ago
Posts: 293
 

I'm the lead data engineer at a mid-sized logistics company, running a mixed stack of Airflow, Snowflake, and various ingestion tools, and I've operated both cloud and on-premise data pipeline platforms for the last five years. Our current production analytics load is served by a hybrid model, using a managed cloud service for core transformations while a legacy on-premise system handles sensitive cargo routing data.

1. **Total Cost of Ownership Calculation**
The cloud list price is easy, but the on-premise bill is deceptive. At my last shop, the all-in cloud subscription was about $28k per month. The on-premise license quote was $14k, but that required three dedicated infrastructure engineers at a fully loaded cost of $45k each per month to manage the Cassandra tuning, RabbitMQ clustering, and control plane patching OP mentioned. The real cost was 4-8 times higher than the license alone.

2. **Infrastructure and Dependency Rigidity**
The "single-binary" promise often assumes a perfect, version-locked environment. Our attempted POC required Cassandra 3.11.9 exactly, which is EOL and has known CVE vulnerabilities. The kernel module for their "lightweight" agent only compiled reliably on RHEL 8.4 with kernel 4.18.0-305, creating a massive security debt and blocking OS updates. The cloud version abstracts all this; the on-premise version makes it your permanent problem.

3. **Operational Overhead and Staffing**
Cloud is a shared operational model. On-premise, as OP discovered, is a full transfer of operational risk. We measured it: the cloud service had a mean time to recovery (MTTR) of under 10 minutes for platform issues. Our on-premise deployment, with the same vendor, averaged 4 hours to diagnose and resolve due to the layers of custom networking, storage, and dependent services we now owned. This required hiring a dedicated platform reliability engineer, a hidden cost never in the vendor's spreadsheet.

4. **Performance and Scaling Profile**
In the cloud, performance scaling is mostly linear and handled by the vendor. On-premise, performance is bounded by your weakest internal dependency. Our workloads saw a 3-4x latency increase on cold queries in the on-premise setup because we could not replicate their cloud's distributed caching layer. Scaling horizontally required a 72-hour lead time for procurement and deployment of new bare metal nodes, versus 5 minutes in the cloud console.

I would recommend the cloud version for any organization without a large, specialized platform engineering team whose sole focus is this tool. If regulatory air-gapping is an absolute, non-negotiable requirement, then go on-premise but demand the vendor provide a fully containerized deployment, not an Ansible playbook, and a guaranteed SLA for support response times. To make a clean call, tell us your team's headcount dedicated to platform operations and your compliance requirements for data residency.


Data doesn't lie, but folks sometimes do.


   
ReplyQuote
(@hannahg)
Reputable Member
Joined: 3 months ago
Posts: 273
 

You're spot on about the hidden engineering costs. That infrastructure rigidity is what kills you, especially when you need to patch a CVE and the whole house of cards comes down.

We ran into something similar with a design system's "self-hosted" component library. The vendor lock-in wasn't about licenses, but about toolchain versions. Trying to update a single npm dependency for a security fix broke the entire build process because their compiler demanded a specific Webpack version from two years ago.

It feels like some companies treat the on-prem version as a grudging compliance checkbox, not a real product someone has to run and maintain.



   
ReplyQuote