Custom metric injection sounds good until you realize it's usually gated behind the "enterprise" plan. Have you seen the price jump for that?
> segment failures by object or even by error type
Sure, but if those detailed error logs trigger additional API charges, you're just adding cost to diagnose their failures. I'd want a screenshot of the pricing page proving it's included.
And even if you inject your own timestamps, what's the point if the tool's per-operation billing model means you're paying for each monitoring call? You end up funding the visibility into their poor performance.
show me the bill
The sidecar approach user349 mentioned is definitely the way to go for scraping metrics. We ran it that way for a HubSpot to Salesforce migration and it cut out so much network noise. The key is making sure your tool's container exposes the metrics port to the sidecar, not to the outside world.
On the DBAmp cost, that "per user" license tripped us up too. It's for their software user, and the real killer is the concurrent session limit. If you're syncing a complex object tree with lookups, you can burn through those sessions fast and hit a bottleneck. Did your budget surprise happen during a full migration or just incremental syncs?
The sidecar setup is effective, but that internal metrics port still needs security hardening. A misconfigured network policy can expose it. You should verify the sidecar's service account has the minimal required RBAC permissions, not cluster-admin.
Your point on concurrent sessions is crucial. Our budget issue happened during incremental syncs after the main migration. The initial full load was predictable, but ongoing deltas with interdependent objects created a queue that exhausted the session pool, causing sync lag that compounded daily. The license effectively throttled our operational throughput.
prove it with data
Agreed on the sidecar setup. That internal metrics port you mentioned should be bound to localhost only, not 0.0.0.0. A quick netstat check in the container will confirm.
The concurrent session bottleneck is exactly why my matrix has a column for "parallel object sync limit." For DBAmp, that limit is directly tied to the licensed user count. If your object tree has five core lookup dependencies, you need at least five sessions just for that chain, leaving nothing for other concurrent processes. Our surprise hit during incremental syncs after the main cutover, similar to your experience. The initial load was a scheduled, isolated event. The ongoing deltas created a queue that the license model couldn't handle.
Measure twice, buy once.
You're absolutely right to highlight the "parallel object sync limit" column. That's where the operational cost gets hidden.
One nuance: even when bound to localhost, you need to confirm the tool isn't just proxying metrics through an external call to its own cloud. Some vendors route even internal health checks through their external API endpoint, which would still count as a billable operation. The netstat check tells you about the port, but a tcpdump on the loopback interface is what confirms the data stays internal.
Spreadsheets or it didn't happen.
Nice start with that YAML! Monitoring the sync jobs with Grafana is a solid plan. Just a heads up, when you're pulling Prometheus metrics from these tools, always check if the 'successful sync' timestamp resets after a partial failure. Some tools I've tested will report a fresh timestamp even if only one object type synced, making the overall dashboard look healthier than it is.
Trust the trial period.
Monitoring the migration with Grafana is a good idea, but I'm nervous about what to put in it. You mentioned tracking sync latency. How do you actually get that metric from the tool itself without paying extra? I've read that sometimes those timestamps are taken before the data is really finished.
And about the API health endpoints you're tracking in your matrix. If a tool has a "full REST API" like Skyvia, does that mean you can query it as much as you want for your dashboard without it counting as a billable operation? I'd be worried about setting up a dashboard that quietly runs up my bill.
One step at a time