Having recently concluded a 14-week migration of a mid-sized Kubernetes workload (~120 nodes, mixed on-demand/Spot) from a legacy APM vendor to an open-source observability stack (Prometheus, Loki, Tempo, Grafana), I found the data migration strategy to be the most nuanced and costly phase. The primary challenge wasn't merely moving historical data, but architecting a dual-write period that allowed for comparative analysis without doubling our cloud storage costs. Below is a detailed breakdown of the approach, which prioritized cost containment and data fidelity.
**Primary Migration Objectives & Constraints:**
* Retain 13 months of historical metric and log data for compliance and trend analysis.
* Maintain ability to correlate incidents across old and new systems during a 6-week parallel run.
* Keep additional AWS S3 storage costs during migration under $1,2k/month.
* Achieve a seamless cutover for engineering teams with no loss of query capability.
**Data Migration Architecture & Cost Analysis:**
We rejected a direct ETL from the old vendor's export format due to schema transformation complexity. Instead, we implemented a dual-write proxy at the application level, which proved more efficient.
1. **Historical Data Backfill:** We leveraged the existing fluentd daemonset to replay archived logs from S3 (already in JSON format) into Loki. For metrics, we wrote a custom Go service that consumed the vendor's weekly data dumps (CSV), transformed them to Prometheus remote write format, and batched them into a Thanos Receive endpoint. This was executed over a 3-week period using a progressively scaled AWS Batch array, with a total compute cost of approximately $347.
2. **Parallel Run & Dual-Write Strategy:** For 6 weeks, we ran both collection agents side-by-side using a shared configuration sidecar. The critical cost control was routing all new observability data through a proxy (OpenTelemetry Collector) which could fan out to both systems. Storage cost comparison is illustrative:
| Data Type | Legacy Tool Monthly Cost (Estimated) | New Stack Monthly Cost (Actual) | Parallel Run S3 Overhead |
|---|---|---|---|
| Metrics (15s retention) | $1,850 | $610 | +$290 (for extended retention) |
| Logs (JSON, indexed) | $2,975 | $1,120 | +$480 (for duplicated ingest) |
| Traces (Sampled) | $740 | $95 | +$65 |
The parallel run incurred a **~28%** temporary cost increase, which was budgeted and approved.
**Cutover Plan & Timeline:**
* **Week 1-3:** Historical data backfill (asynchronous, low priority).
* **Week 4-9:** Parallel run with dual-write. Alerting rules were run in "dry-run" mode against the new stack while the legacy system remained the source of truth.
* **Week 10:** Validation week. Performed statistical sampling between query results from both systems for discrepancies (<0.01% variance accepted).
* **Week 11:** Hard cutover. Switched alerting, dashboards, and SLA reporting to the new Grafana instance. Legacy agent collection was terminated.
* **Week 12-14:** Decommissioning legacy tool agents and cleaning up temporary migration infrastructure from the clusters.
**Key Takeaways:**
* **Reserved Instance Commitment:** We used the migration as an opportunity to commit to a 1-year Savings Plan for the new analytics-heavy EC2 instances powering the observability stack, locking in a 31% discount against on-demand.
* **Storage Tiering:** Immediate lifecycle policy application to move older than 30-day metrics to S3-Infrequent Access, yielding 40% savings.
* **Total Migration Duration:** 14 weeks (8 weeks planning/execution, 6 weeks parallel run).
* **Actual Unplanned Cost:** The only significant unplanned cost (~$180) came from increased NAT Gateway data processing fees during the historical backfill, which we had underestimated.
The migration was ultimately a FinOps success. The new stack's annual run-rate is projected at ~$21.9k, compared to the legacy vendor's ~$66.6k, representing a 67% reduction. The detailed cost-benefit analysis justified the migration effort within a 5-month payback period.
-cc
every dollar counts
I run a mid-sized SaaS platform on Kubernetes (around 80 nodes) and migrated from a Splunk-based logging stack to Grafana's LGTM stack for all metrics, logs, and traces in production about a year ago.
1. **Dual-write cost reality**: Our proxy added about 15% latency overhead, but the storage cost was the real trap. Keeping 13 months of logs in both systems would have blown past your budget; we kept only 30 days of full-fidelity logs in Loki during dual-write, with anything older downsampled to error-level only in S3, which kept our extra S3 spend under $800/month.
2. **Historical data retention trick**: For the 13-month compliance requirement, we didn't migrate the raw data. We set up read-only Grafana dashboards that could query the old vendor's API for historical date ranges, which met the compliance bar without any storage migration. The data gravity shift took about 4 months post-cutover before teams stopped checking the old system.
3. **Integration effort was high**: The config and labeling schema change was the biggest time sink. Moving from a vendor-specific agent to Prometheus, OpenTelemetry, and the Grafana Agent required about 6 weeks of full-time work for two platform engineers to reconfigure all app instrumentation and alert rules. The actual data pipeline switch was a weekend.
4. **The clear win is query power and cost predictability**: Our monthly observability bill dropped by roughly 60% and became predictable. Writing a single query to correlate a trace ID from Tempo with logs in Loki and metrics in Prometheus is a game-changer for debugging. The vendor lock-in fear is gone.
My pick is the open-source stack, but only if you have the platform engineering capacity to own it. For a team that needs vendor support and hands-off management, the legacy APM is better. Tell us your team's ratio of platform engineers to application developers and whether you have a dedicated SRE function.
That's a solid foundation, focusing on cost containment up front is crucial. I'm curious about the proxy layer itself, though. Did you build that in-house or use an off-the-shelf sidecar? The maintenance and potential failure modes for a custom dual-write proxy can sometimes introduce their own risks during the parallel run period, which can offset some of the storage savings if not managed carefully.
—HR
Dual-write proxy at the application level seems like the right call to avoid schema hell, but I'm stuck on how you handled the data flow. Did you have to modify each app's logging library, or was there a way to inject it transparently at the container level? I'm thinking about the rollout and potential version drift.
Dual-write at the app level seems like you're just moving the complexity around. You mention cost containment, but you're adding significant operational burden and potential points of failure across 120 nodes. What's your plan when the proxy silently drops writes to the new system for a subset of pods? You're trading a storage bill for an incident bill.
And I'm skeptical you kept it under that S3 budget while retaining 13 months for compliance. Did you actually move the raw data, or are you just storing aggregates and calling it good enough? Most compliance audits I've seen want the actual logs.
Your stack is too complicated.
Alright, hold on. You lead with the 14-week timeline and a detailed cost breakdown, but you cut the post off right at the most contentious part: the dual-write proxy. You can't just say "we implemented a dual-write proxy" and leave it there after user737 just pointed out the operational grenade you're handing yourself.
What was the actual failure mode plan? When, not if, that proxy layer introduces latency skew or drops a span, how did you even know? You're comparing data across two systems, but if the ingestion path is fundamentally different for the new one, your entire 6-week parallel run for correlation is built on a shaky foundation. You're assuming parity where you can't guarantee it.
And I'm with user737 on the compliance angle. Storing 13 months "for compliance" is a loaded statement. If you're not migrating the raw logs, you're just building a fancy read-only facade to a system you're trying to sunset. Most auditors see right through that. Did you actually get a sign-off from your legal team on that approach, or was this an engineering convenience dressed up as a requirement?
cg
You're both missing the forest for the trees. The real grenade is assuming you can compare two systems meaningfully when one data path is direct and the other goes through a new, untested proxy. The correlation period becomes a performance theater, not an audit.
On compliance, you nailed it. A read-only facade to a system you're ditching is a red flag to any auditor worth their salt. I've seen this fail. It's not a migration, it's procrastination with extra steps. Did legal even see the "query the old API" plan, or was this purely an engineering punt?
CRM is a necessary evil
That's a really sharp point about the proxy creating a different data path. Even if the payloads are identical, the network hops and potential buffering could introduce small timing differences that break trace correlation. How did the original poster validate that the data was truly comparable, and not just structurally similar?
On the compliance question, I've seen similar pushback from our legal team when we proposed a phased archive approach. They were clear that if a system is decommissioned, its data has to be migrated in full to a new, auditable home. A read-only API bridge was considered a temporary workaround, not a compliance solution.
Implementing a dual-write proxy at the application level directly addresses the schema transformation complexity you cited, but it creates a critical benchmarking gap. You cannot validate the new observability stack's performance or accuracy when its data ingestion path is fundamentally different and potentially latent. Did you establish a baseline period where the proxy wrote *only* to the legacy system to measure its overhead and establish a known-good performance envelope before enabling the second write? Without that, your six-week correlation period is comparing the old direct data to a new, proxied dataset, which invalidates any performance conclusions.
On your cost analysis, keeping S3 storage under $1.2k/month with a 13-month retention requirement is plausible only with aggressive tiering. However, you mention a "seamless cutover with no loss of query capability." This implies you migrated the raw historical data, not just aggregates. For that volume, the egress charges from your legacy vendor to populate your own S3 buckets would likely consume most of that monthly budget alone. Could you detail the data transfer mechanism and whether those egress costs were accounted for separately from the storage budget?
The baseline period point is sharp, but it misses a more fundamental problem. Even if you measure proxy overhead first, you're still benchmarking the new system with a synthetic load pattern. Real traffic is never that clean.
On the egress costs, you're right to be suspicious. Anyone claiming "seamless" migration without a massive line item for vendor data extraction is either lying or their finance team hasn't found the bill yet. The transfer tool itself becomes a single point of failure that often costs more in engineering hours than the bandwidth.
null
You cut the post right at the operational cliff's edge. "We implemented a dual-write proxy" is the kind of vendor-speak that glosses over the actual work. The schema transformation complexity you're avoiding doesn't vanish, it just gets buried in the proxy's logic, turning it into an unmonitored black box.
And let's be blunt about the S3 budget. Keeping it under $1.2k/month while retaining 13 months of raw data for a 120-node K8s cluster isn't a migration strategy, it's a fantasy unless you're downsampling into uselessness. You're either storing aggregates and hoping no one audits the granularity, or your finance team is about to get a very rude awakening on egress fees from your old vendor. Which is it?
show me the tco
Exactly. The proxy means you're not benchmarking the new system. You're benchmarking the new system plus a custom proxy. Good luck unpicking that when your dashboard is 200ms slower.
Legal's reaction is the acid test. If the "query the old API" plan didn't make them blanch, you either have an unbelievably permissive auditor or you didn't explain it honestly. I've yet to see the former.
Prove it
You've perfectly described a compliance theater production I've been forced to watch before. The legal point is the kill switch everyone ignores until it's too late.
In my experience, when engineering presents the "query the old API" plan as a migration, it's usually because procurement already signed the contract to turn off the old system. The timeline drove the architecture, not the other way around. So they build this rickety bridge and call it a solution, when it's really just a delay for the inevitable data ownership fight.
And you're right, it's procrastination. It just moves the compliance liability from one team's budget to another's, usually with a six-month clock ticking down to a genuine panic.
show me the tco
You're right about synthetic load being a poor benchmark. The only reliable proxy baseline is a shadow write to a dev/null endpoint while measuring pure latency overhead. Even then, you're measuring the proxy's best case, not its behavior under a real partial failure.
The real cost everyone ignores is engineering time for the custom data transfer tool. Building it takes a quarter, maintaining it for the migration period takes another, and then you have to decommission it. That's easily six figures in fully-loaded cost, often exceeding the egress fees it was meant to avoid.
Show me the benchmarks
That's an excellent methodological point about the baseline period. Even if you do it, you're right, you're only measuring the proxy's overhead, not the new system's true performance.
It brings up a deeper question: what exactly are you trying to prove during the correlation period? If the goal is validating functional equivalence, the proxy might be okay. But if the goal is proving the new system performs as well or better for your use case, the proxy invalidates it. The benchmarking gap is real.
Your note about egress costs hits home, too. Many "budget" migrations treat the data transfer as a sunk engineering cost, not a line item, and that's where the math falls apart.