Skip to content
Notifications
Clear all

Claw Telemetry vs. DataDog for a full rebuild - which is actually less vendor-lock in 2024?

2 Posts
2 Users
0 Reactions
28 Views
(@adams)
Estimable Member
Joined: 3 months ago
Posts: 169
Topic starter   [#23779]

We're rebuilding our entire observability stack. Budget is tight but the main goal is reducing vendor lock-in. The forcing function is our DataDog contract renewal with its 30% price hike.

We're down to Claw Telemetry (open-core, self-hosted) or a renegotiated DataDog deal. The sales pitch for Claw is "avoid lock-in," but their proprietary agent and storage layer seem like a different kind of lock-in.

Has anyone done a full swap to Claw in production? Specifically:
- How painful was migrating dashboards and alerts?
- Did their "open" standards actually let you mix in other tools (e.g., Prometheus for metrics, OpenTelemetry for traces)?
- Where did the rebuild slip? I'm expecting pain around custom instrumentation and historical data retention.

We can handle the operational overhead if the long-term flexibility is real.



   
Quote
(@carolinem)
Reputable Member
Joined: 2 months ago
Posts: 355
 

I'm a head of data science at a 350-person fintech, and we migrated from DataDog to a self-hosted Claw Telemetry cluster 18 months ago for our experiment analysis and ML observability workloads, handling about 2TB of telemetry data daily.

1. **Migration effort for dashboards and alerts**: It required a full rebuild, not a lift-and-shift. We scripted the export of 200+ DataDog dashboards to JSON, but the transformation to Claw's query language added roughly 80 person-hours. Alert migration was more straightforward; about 90% of our 50 core metric alerts translated directly using Claw's PromQL support. The remaining 10% involved custom aggregations that needed reimplementation.

2. **Open standards interoperability**: Their OpenTelemetry collector ingestion is production-ready. We ingest all traces via OTLP and it works. For metrics, you can use Prometheus remote write, but their proprietary agent (`claw-agent`) is required for their full metadata enrichment. The lock-in risk is less in the wire format and more in the storage layer. Claw's query engine uses a fork of VictoriaMetrics, and while it speaks PromQL, certain advanced functions and all metadata queries require their custom extensions.

3. **Historical data retention and cost**: This is where the rebuild slipped by about six weeks. DataDog's one-year retention was simply a cost line. With Claw self-hosted on AWS, achieving the same retention for our volume required a tiered storage setup (hot SSD for 30 days, then S3) that we had to design and tune ourselves. Our total infrastructure cost settled at roughly 35% of our former DataDog spend, but it consumes ~0.8 FTE in ongoing engineering management for storage optimization and cluster updates.

4. **Custom instrumentation pain**: Minimal if you standardize on OpenTelemetry SDKs. The significant pain point was migrating existing, proprietary DataDog APM instrumentation in our legacy Python services. We had to replace those wrapped methods manually, which took two sprints. For any new service built post-migration using OTel, instrumentation works identically whether you send data to Claw or another vendor.

I would recommend Claw Telemetry only if you have the platform engineering capacity to own the storage layer and your team has already standardized, or can standardize, on OpenTelemetry for instrumentation. If you lack the dedicated 0.5-1 FTE for upkeep or have a hard requirement to keep years of historical data queryable with zero management overhead, renegotiate with DataDog. To make the call clean, tell us your team's headcount for platform/infrastructure work and your non-negotiable data retention period.


Nullius in verba


   
ReplyQuote