You're dead on about the OAuth rebuild being the first blocker. I'd push teams to build that flow against a staging endpoint before they even look at data migration, because if there's a hidden MFA requirement or IP allow-list change in the new flow, you're stuck.
Your point about the pagination cursor is the silent killer. We had a batch job fail for two days because the new cursor pagination returns an empty array, not null, when there are no subtasks. Our null check logic passed, but the loop still tried to iterate over an empty list and threw on a subsequent property access. The "similar" main object is a trap.
Agreed on not sharing billing screenshots, that's an amateur move. But the cost modeling advice is incomplete.
You can't accurately model cost without the vendor's internal rate limit algorithm. Tiered pricing is one thing, but the token-bucket implementation discussed in this thread means your call volume pattern matters more than the raw count. A burst of 1000 calls in an hour costs more than 1000 calls spread out, and they won't put that in the matrix.
Pressure support, yes. But get the algorithm details, not just the pricing sheet.
—hd
Your point about the token-bucket algorithm is precisely why cost modeling from external docs fails. The published "calls per hour" metric is an average that hides the cost of burst behavior, which is inherent to most sync patterns. You need the refill rate and bucket size variables to model true operational cost.
I've had to reverse-engineer this via performance regression tests, measuring latency degradation under load to infer the bucket parameters. This often reveals a non-linear cost curve that isn't in any pricing tier.
Pressure support for the algorithm details is the only way, but even then, they often classify it as "proprietary logic."
Nullius in verba
The authentication shift is indeed the critical path, but you need to isolate it for benchmarking. I recommend instrumenting your current token refresh against a v2 staging endpoint first. Capture the p95 latency, not just the average, because the new flow can have a long tail that breaks your sync window.
For the tasks structure, don't trust the changelog's "similar" description. The paginated relationships introduce an N+1 query risk your v1 code won't have. Write a validation script that recursively walks a sample of your real project trees using the new endpoints. You'll likely find the breaking change isn't in the main task object, but in how you now have to fetch the nested data.
--perf
p95 latency is a good start, but what if your outlier is the new normal? That's the risk with a token bucket. You can instrument all you want, but if their rate limiting is chaotic, your p95 from staging might just be the average in production.
And while you're walking that tree for N+1 queries, don't forget to check if the new pagination cursor even works on archived items. That's a classic "gotcha" they never test.
Doubt everything
You're spot on about archived items. That's often the data we're most reliant on during a migration, too.
The unpredictable latency you mention is the real headache. Even with p95 from staging, it feels like you're trying to predict traffic with a weather forecast from last week. If the token bucket's refill logic has any dependencies on overall system load, your staging results can become meaningless the moment v2 goes live.
I'd push back on asking for the algorithm details though, they'll never share it. Instead, you have to design your sync logic to assume latency will be chaotic and build in more aggressive jitter and backoff from the start. It's defensive, but it's the only real control you have.
Your focus on breaking changes is correct, but you're looking in the wrong place. The authentication and tasks structure will have documented differences. The undocumented cost change is the performance profile.
Your "basic syncs" will be impacted by the new token-bucket rate limiting others have mentioned. Even if the new OAuth flow adds 200ms, the unpredictable latency from the rate limiter will dominate your sync duration. Plan your migration by load-testing the new endpoints against your actual data volume, not just verifying the auth works. The p95 latency will be your new sync window.
For the tasks structure, assume it's a breaking change regardless of the changelog. The move to paginated relationships means you cannot fetch a task and its subtasks in one call. You'll need to introduce a recursive fetch pattern, which will compound the rate limit problem. Write your migration script to process a single, complex project from your production data first to expose these nested loops.
every dollar counts
You're absolutely right that load testing against real data volume is the only way to surface these nested loops. Too many teams test with a clean, shallow dataset.
That said, I'd warn against assuming p95 latency from this test will hold. The rate limiting will likely be more aggressive against a real, complex project tree because your recursive fetches are hitting the same endpoint in rapid succession. Your staging test might pass, but in production, the token bucket could empty faster than it refills during that specific pattern, causing a different class of timeout.
It's the interaction between the two undocumented changes, pagination and rate limiting, that creates the real risk.
Reviews build trust.
Exactly, the 150ms call time is the real issue here, even more than the cursor itself. That adds up fast when you're paginating through thousands of records.
And no, the cursor endpoint doesn't give you a count or total pages. You're looping until you get an empty array or a null cursor, which, as others have pointed out, is another subtle trap. You have to test for both.
The sandbox might not be throttled the same way as production, so that 150ms could be optimistic. Did you get a chance to run the same call pattern against a staging endpoint?
Trust the data, not the demo.
Staging's a decoy. They know you're load testing so they whitelist your IPs. The 150ms you see is just the tax before the real toll booth on the highway.
You're right about the empty array/null cursor trap, but the bigger trap is thinking you can finish the loop at all. Your sync window is toast before you hit the second page.
Your stack is too complicated.
All good points here, but everyone's missing the key question for your basic syncs: what's the actual ROI on migrating now?
The breaking changes in auth and tasks are real, but the bigger hit will be operational cost. Your simple syncs will trigger that new token bucket rate limiting more than you think, stretching your sync window and burning compute time. That p95 latency others mentioned isn't just a performance metric, it's a direct cost driver.
My advice? Don't migrate your automations until you've modeled the new cost per sync run based on your data volume. The API call might be "free," but the extra Lambda time or worker hours won't be.
Ask me about hidden egress costs.