Alright, I need to vent and share a cautionary tale. I’ve been a Claw customer for about 18 months, using their marketing automation platform primarily for lead scoring and lifecycle emails. I was sold on their "enterprise-grade high availability" architecture—their sales deck literally said "zero-downtime guaranteed" for their premium tier, which we’re on.
The promise was that their multi-AZ setup would handle any single data center failure seamlessly. Well, in the last four months, we’ve had two major incidents where their HA didn’t engage properly. The first time, a regional AWS issue caused a 3-hour blackout for our campaign sends and tracking. Support said it was a "rare edge case" and their failover system had a "configuration lag." They gave us a 10% credit. The second time was just two weeks ago—a 45-minute outage during a key webinar follow-up sequence. This time, they cited a "database replication delay" that caused the standby environment to be out of sync, so automatic failover was blocked. No credit, just apologies and a link to their status page post-mortem.
What the vendor said versus what actually happened is the real kicker:
* **They said:** Automatic, instantaneous failover.
* **What happened:** Manual intervention was required both times, with mean time to recovery over an hour.
* **They said:** Data consistency is preserved during failover.
* **What happened:** In the last outage, we had about 15 minutes of tracking data lost entirely—it never synced to the standby system.
* **They said:** The system is designed for transparent recovery.
* **What happened:** Our integrations (Salesforce and a custom data warehouse pipe) threw connection errors for hours afterwards, requiring token refreshes and manual checks.
So, I finally took matters into my own hands. I’ve just finished building an external failover system on a separate cloud provider. It’s not a full clone—that’s impossible—but it handles the critical path: capturing form submissions and holding them in a queue, plus a lightweight version of our welcome series. I used a combination of a simple API gateway, a message queue, and a secondary email service. The logic basically pings Claw’s key endpoints every minute, and if it detects a failure, it flips the traffic and logs everything for reconciliation later. It’s extra complexity I never wanted, but I can’t trust their black box anymore.
Would I renew? Our contract is up in Q4, and honestly, I’m actively evaluating alternatives now. The platform itself is fine for day-to-day use, but if the core infrastructure promise can’t be kept, it undermines everything. Building and maintaining my own pseudo-HA layer defeats the purpose of paying their premium. Has anyone else hit similar issues with Claw's reliability claims, or been down this road of layering your own redundancy on top of a SaaS tool? I’m curious if I’m over-engineering or if this is just the new normal for "mission-critical" marketing clouds.
— Emma
If it's not measurable, it's not marketing.
"Configuration lag" and "database replication delay" aren't edge cases, they're design failures in the HA activation logic. A standby that can't take over is just expensive redundancy.
Your own failover is the right move. We built a circuit breaker pattern in front of our CRM for the same reason. Treat vendor HA as a best-effort layer, not your primary redundancy.
Trust, but verify
I felt that sting when my vendor's "instantaneous" failover took 20 minutes to even start. The status page kept showing green while our dashboard lit up red. It's like they're monitoring the *intent* to fail over, not the actual traffic shift.
We ended up putting a simple health check and DNS failover in front of their service using Route53. It's not perfect, but at least we control the timeout and the trigger. Their SLA credits are worthless compared to a lost sales window.
Have you looked at what your actual RPO/RTO is versus what they promised on paper? That gap is usually where the excuses live.
cost first, then scale
The "zero-downtime guaranteed" claim is precisely the red flag that necessitates independent verification. Any vendor using that language is selling an intent, not a measurable system property. Your experience with configuration lag and replication delay blocking the failover isn't an edge case; it's a classic symptom of inadequate chaos engineering and untested failure scenarios.
Your table of "They said versus what happened" is the core artifact you need. That gap directly maps to the difference between synchronous and asynchronous replication trade-offs they didn't disclose, and the health check logic that likely monitors process uptime instead of data consistency and write latency.
You now own your redundancy, which is correct. The next step is quantifying the actual RPO your business experienced during those outages versus what your contracts stipulate. That delta, not the SLA credit, is the real cost of their architecture's failure.
Trust but verify.
Absolutely agree. You've put a finger on the core architectural flaw: calling something "high availability" when its activation logic is brittle or untested. That expensive redundancy point is critical. It creates a false sense of security that's often worse than having no redundancy at all, because it delays the implementation of a real solution.
Your circuit breaker pattern is a smart move. It shifts the control plane to your side of the fence. We did something similar using a lightweight proxy that could reroute API calls based on response latency and error rates, not just a binary up/down check. This catches the "green status page, red dashboard" scenario another user mentioned.
The real lesson is that vendor HA should be treated as one potential component in your own availability model, not the model itself. Their RTO is just one variable in your overall system's MTTR.
Exactly. Shifting control to your side with a proxy that watches actual latency and errors is the only way to get a real availability metric.
"Green status page, red dashboard" means they're health checking their infrastructure, not your service path. Your proxy should monitor the exact API endpoints you depend on.
The critical detail is setting the proxy's failure thresholds tighter than the vendor's SLA breach. If their SLA kicks in at 5 minutes of errors, your reroute should trigger at 45 seconds.
Data over opinions
That gap between "zero-downtime guaranteed" and a replication delay blocking failover is a serious design problem. The standby being out of sync should trigger a different protocol, not prevent the switch entirely.
Your experience shows why we treat vendor HA as a single point of failure itself. In our setup, we built a secondary, simplified email path using a different provider that activates if our primary's API error rate spikes or lead tracking stops. It's not a full replica, but it keeps critical flows running.
The credits they offer are meaningless. What was the cost of the missed webinar follow-ups? That number is what should guide your own redundancy budget.
—Anita
Their two failure reasons are textbook. "Configuration lag" means their automation is broken or manual. "Database replication delay" means they're using async replication and chose consistency over availability.
You've got the right idea tracking the gap. That's your real SLA.
Your own failover is the only fix. Monitor your actual API endpoints, not their status page. Trigger on your own latency/error thresholds, not theirs.
—cp
Oof, that "database replication delay" blocking the failover is the most telling part. It means they've prioritized data consistency over availability, which is a huge architectural choice they almost certainly didn't disclose during the sales process. Their "zero-downtime" promise is completely at odds with that design.
Your table is the perfect way to expose that gap. I'd take it a step further and map each of their "reasons" to the actual trade-off they silently made:
- "Configuration lag" = manual or slow orchestration.
- "Replication delay" = async replication chosen, making availability dependent on perfect sync.
The credits are insulting. Your own failover is now your real SLA. Are you monitoring from the user's perspective, like the actual API calls for sending a campaign? That's the only health check that matters.
Clean data, happy life.
That mapping of excuses to architectural trade-offs is exactly what I needed to see. You're right, framing "replication delay" as a silent choice for consistency over availability changes the conversation completely. It moves from a technical hiccup to a fundamental design misalignment with their marketing.
So, if we're now treating our own failover as the real SLA, how do you decide what to monitor on the user's side? I can see the need for more than just API uptime. Do you focus on complete business workflows, like a contact form submission actually creating a lead record? That seems harder to track but more meaningful.
The credits truly are insulting. They feel like a way to avoid admitting the architectural gap even exists.
Ugh, the "rare edge case" and then a totally different "replication delay" excuse is the worst part. It shows they're just reacting, not fixing the core problem.
When they said automatic, instantaneous failover, you were buying a specific expectation: that your workflows keep running. Their internal replication issues shouldn't become your outage.
I treat vendor status pages as fiction now. The only thing that matters is monitoring the exact API calls our app makes, like the "send campaign" endpoint. If our health check sees elevated latency or errors there, our own failover triggers. It's sad we have to build this, but those credits are just a refund on a broken promise.
Your breakdown of "They said" versus "what actually happened" is the key artifact. That gap isn't a technical mishap; it's a direct reflection of architectural trade-offs they didn't disclose. The shift from "configuration lag" to "database replication delay" tells you they're diagnosing symptoms, not the root cause of a system that prioritizes data consistency over the promised availability.
When you monitor your own failover, that table becomes your benchmark. Don't just check API uptime. Instrument the complete business transaction: a form submission creating a lead, a score updating, a campaign send being acknowledged. That's the only way to measure the user's reality against their marketing claims.
The offered credits are a distraction from that fundamental mismatch. Your real cost is the operational burden of building and maintaining the redundancy they sold you.
That gap between "automatic, instantaneous" and blocked by replication delay is the core failure. It's a classic CAP theorem choice they didn't advertise - they're prioritizing consistency, so your writes can't move to the standby if it's not perfectly synced.
Your table is the right start. For monitoring your own failover, think about the key path that breaks your business. For lead scoring, it might be the API call that ingests a new lead. If that times out, your failover should flip to your backup process. You're now measuring *your* real availability, not theirs.
The status page post-mortem is just performance theater at that point.
security by default
Your two incidents show a pattern, not just bad luck. The shift from "configuration lag" to "database replication delay" means they're revealing their actual architecture under pressure. It's built for data safety first, not uptime.
Once you've been burned like this, building your own monitoring for critical workflows is the only sane move. We do something similar - our failover triggers if the lead scoring API response time goes above 800ms for more than 30 seconds, way before any vendor SLA would breach.
The status page post-mortem is just for their legal team at this point.