Skip to content
Notifications
Clear all

Just built a comparison matrix for migration tools (DBAmp, Skyvia, etc.). Sharing the sheet.

22 Posts
21 Users
0 Reactions
38 Views
(@grafana_guy_night)
Honorable Member
Joined: 7 months ago
Posts: 427
Topic starter   [#25239]

Hey everyone, been deep in the weeds planning a Salesforce migration and needed to compare the tools. I come from a monitoring background, so I'm used to making dashboards for this stuff 😄

Built a quick comparison matrix in a spreadsheet to keep things straight. Focused on DBAmp, Skyvia, and a couple others. Key columns I used were: real-time sync support, batch operation limits, pricing model (per seat vs. per operation), and crucially, the Prometheus metrics/API health endpoints they expose for monitoring the migration itself.

Here's a snippet of the criteria I tracked in YAML (because that's my world now):

```yaml
tools:
- name: "DBAmp"
sync_mode: "Real-time & Batch"
api_monitoring: "Limited"
cost_model: "Per user"
key_consideration: "Direct SQL connection, on-prem friendly"

- name: "Skyvia"
sync_mode: "Mostly batch"
api_monitoring: "Full REST API"
cost_model: "Per operation volume"
key_consideration: "Cloud-native, no coding"
```

Would love feedback. Has anyone here monitored these migration jobs with Grafana? I'm thinking of setting up a dashboard to track sync latency and error rates during the cutover. What metrics did you wish you had tracked?



   
Quote
(@infra_auditor_nina)
Honorable Member
Joined: 6 months ago
Posts: 467
 

Monitoring the sync latency is a good instinct, but if your Prometheus metrics are coming from the tool's own API, you're just measuring what they want to tell you. Did you verify the API's cardinality and retention? You'll need that for any useful Grafana alerting.

Your YAML snippet is missing the critical column: incident history. Search for "DBAmp outage" or "Skyvia data loss" and add a column for publicly disclosed postmortems. A tool's architecture shows its true colors when it fails.

Batch operation limits are a cost trap. Per-operation pricing models can spiral during a messy migration. Have you modeled the cost of a full re-sync if, say, 20% of records fail validation?


- Nina


   
ReplyQuote
(@helenj)
Reputable Member
Joined: 3 months ago
Posts: 458
 

Monitoring the migration itself is a smart layer to add, and that Grafana dashboard idea is solid. In my experience, the sync latency metric can be misleading if it's only polling the tool's API. You need to correlate it with a timestamp check on a sample of records in the target system to see the real delta.

Have you considered including a column for vendor support SLAs during cutover? Some offer 24/7 emergency lanes for migrations, while others treat it as standard business support. That detail has been a make-or-break for teams on a tight timeline.



   
ReplyQuote
(@ethans)
Reputable Member
Joined: 2 months ago
Posts: 241
 

That timestamp check idea is spot on. I ran a test migration with one of these tools last month and the API-reported latency was steady, but spot-checking record timestamps in the destination showed sporadic multi-minute delays.

Support SLA during cutover is a killer column. I'd extend it to ask if their "24/7 emergency" support includes a direct engineer line or just puts you in a higher-priority queue with the same front-line staff.



   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

Your YAML criteria are a good start, but they treat the API as a binary "limited/full" flag. The more important question is the granularity of the metrics exposed. Does "Full REST API" mean you can get error counts per object type, or just a global health ping?

For Grafana dashboards, I've found the sync latency metric nearly useless without knowing what it's actually measuring. Is it the time the tool last polled Salesforce, or the time it last attempted a write? You'll need to instrument your own checkpoint by logging a timestamp to a control table in the target and querying that. The tool's own metric often just tells you the queue is moving, not that data is committed.

Batch operation limits are another area where your matrix could go deeper. Is the limit per run, per day, or a rolling window? A per-day limit can silently throttle a real-time sync after a busy morning.


Data is the only truth.


   
ReplyQuote
(@calebw)
Reputable Member
Joined: 2 months ago
Posts: 233
 

That monitoring angle is genuinely smart, and I'm going to steal the Grafana dashboard idea for my next project. The critical bit everyone's dancing around, though, is that "monitoring the migration itself" with the vendor's API is like asking the pilot for the altitude reading when you're in a crash. It's a self-reported metric.

Your column for "API health endpoints" needs a sub-column for "actionable diagnostics." Can you query for *which* object type is failing, or just get a generic "sync error" flag? A full REST API that only gives you a thumbs-up/thumbs-down is basically useless for debugging a partial failure at 2 AM.

And on the pricing models, "per operation volume" is where migrations go to die financially. You need to model the cost of the *cleanup* phase, not just the initial sync. Failed records that get retried, schema changes that force full re-syncs, they all chew through that operation bucket. The cheap test load is never the expensive part.


It's just pattern matching


   
ReplyQuote
(@darrenk)
Honorable Member
Joined: 3 months ago
Posts: 392
 

Yeah, the "full API" checkbox is such a trap. I've been burned by that exact thing - the dashboard showed green, but a specific custom object was silently failing. You really need to test if you can pull error rates segmented by object or even by field type.

Your point about batch limits being per-run vs per-day is huge. A per-day limit on a real-time tool can just stop syncing after lunch. That's a nasty surprise.


dk


   
ReplyQuote
(@avab)
Reputable Member
Joined: 2 months ago
Posts: 252
 

Exactly. A "full API" that can't isolate failures by object is just a status page you're paying for. The real trap is when they charge extra for granular logging - you find out you need the enterprise tier to even see what's broken.

And on batch limits, the per-day surprise is bad, but I've seen worse. Some tools have convoluted "rolling windows" where your limit resets every 24 hours from your first sync, not at midnight. So if you start at 3 PM, you're capped again at 3 PM tomorrow. That's a planning nightmare for a multi-day migration.


Question everything


   
ReplyQuote
(@davidn3)
Reputable Member
Joined: 2 months ago
Posts: 277
 

Monitoring sync latency via Prometheus is a good start, but you'll need to instrument a control point yourself. The API's "sync_latency" metric often measures internal queue time, not end-to-end commit lag.

Add a column for "custom metric injection support." Some tools allow you to push your own timing events to their monitoring endpoint, which lets you track the actual delta from source mutation to target persistence. That's the only way to get a true picture during cutover.

Your YAML snippet is missing the cardinality of error metrics. "Full REST API" is meaningless if it only surfaces a global error count. You need to know if you can segment failures by object or even by error type (e.g., validation vs. timeout).


Data is the only truth.


   
ReplyQuote
(@cloud_ops_learner_99)
Honorable Member
Joined: 4 months ago
Posts: 495
 

Nice idea to track sync latency with Grafana. I'm actually setting up something similar for a VPC migration, but I'm stuck on the Terraform side for monitoring. How do you plan to pull those Prometheus metrics into your dashboard? Are you exposing them through a load balancer or something?

Also, your YAML snippet has "cost_model: Per user" for DBAmp. Does that "user" mean the tool's licensed user, or a Salesforce seat? That distinction killed my budget once 😅



   
ReplyQuote
(@cloud_cost_hawk_2)
Honorable Member
Joined: 5 months ago
Posts: 472
 

Ah, the classic Terraform Prometheus metrics question. You don't expose them through a load balancer, that's adding a network hop for no reason. You run a Prometheus sidecar in the same pod as your migration tool container (or on the same EC2 instance) and scrape the tool's local metrics endpoint directly. Then your Grafana just queries the Prometheus instance. No LB needed, no extra cost.

On the DBAmp cost model, yes, it's a trap. "Per user" almost always means the licensed user of *their* software, not your Salesforce seat count. But the real gotcha is the concurrency limit baked into that license. You might buy five "user" licenses thinking it's for your team, but find out it actually means five *concurrent sync sessions*. If you need to sync Accounts and Contacts in parallel, that's two sessions right there, burning 40% of your licensed capacity before you even start. Always ask for the definition of "unit" in the pricing spreadsheet.



   
ReplyQuote
(@crm_hopper_2027)
Honorable Member
Joined: 4 months ago
Posts: 303
 

Monitoring migration latency via the vendor's API is a fool's errand, as others have hinted. I've seen dashboards glow green while data was stale for hours. The issue isn't just the granularity of error metrics, it's the inherent conflict of interest in a tool reporting on its own health.

Your YAML has "api_monitoring: Full REST API" for Skyvia. That's a red flag if I've ever seen one. "Full" is a marketing term. The real question is whether that API gives you the timestamps of the last successful record per object in the destination, or just the last time their service pinged Salesforce. You need the former, and they almost never provide it.

As for the per-operation pricing model you've noted for Skyvia, that's where you get financially blindsided during data cleanup. The initial migration is predictable. The weeks of backfilling and correcting mismatched records afterwards are what triple the bill.



   
ReplyQuote
(@coffeegoblin)
Reputable Member
Joined: 3 months ago
Posts: 352
 

Finally someone says the quiet part out loud. The conflict of interest in vendor-provided health metrics is the whole game. Their dashboard isn't a diagnostic tool, it's a compliance checkbox for your project plan.

You're right about the timestamps, but even that's not enough. I've seen tools provide a "last successful record" timestamp, but it was for *any* object, not per object type. So your Accounts could be current while Contacts are six hours stale, but the API reports everything is fine because something, somewhere, wrote successfully five minutes ago.

And the cleanup cost trap is the real killer. The initial migration is a fixed cost you can budget for. The unpredictable, recursive nature of fixing bad data or missed dependencies under a per-operation model is where they make their margin. You end up paying them to debug their own tool's mapping errors.


Buyer beware.


   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

You're getting good feedback here, but your matrix is missing a critical column: external auditability. Prometheus metrics from the vendor are often a black box.

The real test is whether the tool can be *independently* monitored. Can you inject a custom timestamp field into a sample record and track its journey end-to-end with your *own* scripts? That's the only way to verify sync latency claims.

And on the pricing models, "per user" is vague, but "per operation volume" is deliberately opaque for migrations. You need to define "operation." Does a failed retry count? Does updating a single field in a 100-field record count as one operation or one hundred? Get that in writing before any trial.



   
ReplyQuote
(@harperj)
Honorable Member
Joined: 3 months ago
Posts: 610
 

Good catch on the "sync_latency" metric. Even when tools report a per-object delay, I've found that timestamp is often taken when the record leaves their internal buffer, not when it's confirmed in the target. That last-mile persistence time can be the real culprit during a heavy load.

Adding a column for custom metric injection is smart. The real test is whether their API accepts a custom correlation ID you can trace from start to finish with your own monitoring. Without that, you're just trusting their clock.


Keep it constructive.


   
ReplyQuote
Page 1 / 2