Did you see the Datadog incident last week? The one that took down their metrics, APM, and logs for a significant period? It was a stark reminder for me.
We rely on them for everything—monitoring, alerting, dashboards. When their platform hiccups, our entire visibility goes dark. We're flying blind, and our internal teams start asking questions we can't answer. It got me thinking hard about the risk of putting all our observability eggs in one basket.
I've been through a major SaaS migration before (ERP systems), and the principles feel similar. Vendor lock-in isn't just about cost; it's about resilience. I'm starting to seriously evaluate a split-stack approach now, maybe for critical systems. Something like:
* Keeping core infrastructure metrics on one platform
* Routing application logs elsewhere
* Using a third for synthetic checks or alert routing
The trade-offs are real—more management overhead, potential data correlation headaches, and definitely more complex contract negotiations. But after seeing a single point of failure impact so many teams simultaneously, the calculus is changing.
Has anyone else been re-evaluating their observability vendor strategy after this? I'm particularly curious about:
* Practical experiences with multi-vendor setups. Is the complexity nightmare people say it is?
* How you decide which telemetry data to separate first.
* Any lessons from contract negotiations that prevent punitive pricing when you're not sending "all" your data.
- h
Data is sacred.
Vendor splits just give you multiple single points of failure. Now you manage three broken dashboards instead of one.
The real lock-in is your data format. If you own that, you can switch. If you're scared of flying blind, keep a tiny, separate system for basic up/down and latency checks. Something stupid simple that never breaks.
Adding three vendors for logs, metrics, and traces means you'll never actually correlate anything during an incident. That's when you go blind.
Simplicity is the ultimate sophistication
Yeah, that's a solid point about correlation. When things are on fire, you need to see the whole picture, not toggle between three different UIs.
But your note on data format is the real key. If you're logging to a vendor's proprietary schema, you're truly stuck. I think the ideal middle ground is a core platform for correlation *plus* that separate, dead-simple heartbeat system you mentioned. Then you at least have a baseline truth when the fancy stuff fails.
It's not about avoiding lock-in entirely, it's about managing the blast radius.
Let the machines do the grunt work
Totally agree on the data format being the critical lock-in. I've seen teams get absolutely paralyzed during migrations because they can't untangle years of custom log enrichment from a vendor's schema.
Your point about a baseline heartbeat system is spot on. We actually set up a tiny, self-hosted instance of a simple monitoring tool for just this reason. It pings our core endpoints and sends a basic SMS if it ever fails. It's saved us twice now when the "single pane of glass" went completely opaque.
Maybe the trick is designing your primary system to export normalized data to a cheap object store as a backup? That way you keep the correlation power for daily use, but have an escape hatch if the main UI goes down.
Beta tester at heart
The idea of a baseline heartbeat system is crucial. I run a small synthetic check from a separate provider that hits our critical login endpoint. It costs almost nothing and has pinged us during internal network issues the main platform couldn't see.
Your point about exporting normalized data to an object store is the real strategy, though. Datadog's Log Archives and Metric Export features are built exactly for this. It means your primary correlation isn't lost, but you're also not trapped. The lock-in happens when you don't use those exits.
null
That's a smart use of their export features. The operational discipline required to actually test restoring from those archives is often the missing piece, though.
Teams configure Log Archives to S3 but never run a drill to query from it during a simulated vendor outage. Without that practice, you're still flying blind for the first critical hour. It turns a technical escape hatch into a procedural one.
Have you built any automation to periodically validate that your exported data is queryable by a secondary tool?
Commit early, deploy often, but always rollback-ready.
You've put your finger on the operational gap. We run a quarterly drill using a scheduled dbt job that pulls a sample from the archived S3 data, transforms it to match a simple internal schema, and loads it into a test Snowflake table. The job then fires a test query.
The failure mode we've caught isn't the export itself, but the IAM permissions and partition structure silently breaking after a platform update. The automation is simple, but it forces us to maintain the escape path as living infrastructure, not just a checked box.
Yeah, that outage really got my attention too. We just started using Datadog for some basic monitoring, and this incident is a bit of a wake-up call for us newbies.
Your idea of a split-stack is super interesting, but I have a newbie question about the overhead. How do you handle the cost and time of learning and managing three different platforms? Is the extra resilience worth potentially slowing down your team?
That point about procedural discipline is so true. We fell into that exact trap years ago with a different platform. We had the backup pipeline, celebrated the checkmark, and then during a real scramble discovered our query patterns relied on metadata fields that weren't in the export. The escape hatch was there, but we'd forgotten to pack the map.
What saved us later was baking a validation step into our deployment pipeline for any dashboard or alert that used "critical" data. If it depends on Datadog logs, the CI job also has to successfully run a stripped-down version of the query against a week-old sample in our backup store. It's not a full drill, but it constantly validates the data shape is usable.
It adds a few minutes of cycle time, but it means the team's operational knowledge of how to query the data stays current. The automation isn't just for the data, it's for keeping our own skills sharp.
Let's keep it real.
Yeah, that hit home for me too. We're just starting with Datadog and this was the first big vendor outage I've witnessed from the user side.
Your split-stack idea makes sense, but I'm already worried about the correlation problem others mentioned. Maybe starting with a separate, dead-simple system for just a couple core health checks is a more manageable first step? That way you at least have a baseline when the main platform blinks.
Your split-stack evaluation is exactly where more teams should be. The ERP comparison is valid, but the operational tempo is different; an ERP outage might slow month-end closing, but an observability blackout happens during an active incident, which is a force multiplier for chaos.
The correlation headache you mentioned is the primary technical argument against splitting. But that's often a failure in abstraction. We've had success treating the observability layer itself as a consumer, not the source of truth. All telemetry is routed to a durable, vendor-neutral object store first, using OpenTelemetry or a simple fanout. Then Datadog, or a second tool, pulls from that. This adds latency, but it means the correlation happens on data we already own and control. If the primary tool fails, we can spin up a Grafana instance pointed at the same bucket in minutes, rather than waiting for an export to complete.
The real cost isn't in managing three platforms, it's in designing that decoupling layer from day one. If you're already deep in a vendor's ecosystem, that retrofit is painful.
Boring is beautiful
Couldn't agree more on the drill being the difference. We call it the "Sunday morning test" - could you restore visibility before your coffee gets cold?
Our automation is embarrassingly simple: a Lambda function fires weekly, tries to fetch the last hour of logs from the backup bucket via Athena, and pings a Slack channel with a thumbs-up or down. The real cost wasn't the setup, but the time we spent fixing the first three failures - broken partition schema, expired STS tokens, you name it.
The box was checked, but the path was overgrown. Now the team actually trusts the backup.
Cloud costs are not destiny.
Completely agree that vendor splits can just create more failure points to babysit. I've seen teams drown in dashboard fragmentation.
But your point about the data format is the real key. We've shifted to pushing all raw events to S3 first, in OpenTelemetry format, before anything touches a vendor's UI. The dashboards and alerts become just views on top of that owned data. It means the "single point of failure" is our own cheap storage, and the vendor tools become replaceable layers.
The tiny separate system is still essential though. We have a couple of synthetic pings from a cheap provider running to an internal endpoint. That's the thing that tells us, with certainty, if we have a platform problem or a real problem.
Let the machines do the grunt work
You're right about the calculus changing. But the split-stack approach you outlined just gives you three smaller single points of failure. You're still vendor-locked, just in three different places now, with triple the contract headaches.
The real shift isn't about splitting vendors, it's about owning the data pipeline. The comments here about pushing to an object store first using OTel are the key. That becomes your durable source. Let Datadog or any other tool be a disposable consumer of that data. It's more initial work, but it means your correlation happens on data you control, and swapping a UI layer becomes a weekend project, not a migration.
Been there, migrated that
Exactly. The "three smaller single points of failure" argument is a solid one I've seen trip teams up. It can feel like you're just swapping one monster for three gremlins.
Where I've seen this work is when the split is by *failure domain* rather than just function. Don't use three closed platforms. Use something like a Prometheus/Grafana stack for metrics (which you can self-host or run on someone else's metal) and keep your logs in that OTel/S3 lake. Now your "vendors" are your own object storage API and an OSS stack you can forklift. The contracts vanish.
But you're dead right about the core principle: own the pipe. The weekend project angle is real. We did exactly that last year when a vendor's price hike hit, and migrating the visualization layer was mostly just rewriting some Terraform modules. The data was already ours.
Keep deploying!