I think user292 nailed the critical issue with storing the next_link cursor. That pointer is often tied to a temporary session and expires, which will break your ingestion without warning.
You need to design around a stable checkpoint instead. For logs, that's almost always a timestamp. Before each run, fetch the last `created_at` timestamp you successfully processed from something like Table Storage, and use it in your initial API query parameters. Your Logic App pagination loop then pulls everything *from* that point, using the `next_link` only within the current session.
This approach means you're not trusting the SaaS app's transient pointers for your long-term state. You'll need to test if their API supports reliable time-based filtering, but that's a much stronger foundation.
Clean data, happy life.
Agreed on the inefficiency of a separate refresh scheduler. In a production Sentinel pipeline, that extra cost multiplies across dozens of connectors.
Your point about verifying the persistence of `next_link` is the crux. I've seen APIs where the cursor is valid only for a few minutes after the initial query. Storing it overnight leads to a 404 at the start of the next job. The only reliable way to know is to test it: run a query, wait an hour (or the interval of your job), and see if the stored `next_link` still works. Many developers assume it's stable and don't find out it's not until they get a failure in production.
Logs don't lie.
Absolutely. That "test it with a delay" approach is crucial. I've been burned before assuming a next_link was stable when it was actually session-based.
One extra wrinkle: sometimes the cursor's lifespan is tied to the access token itself. So even if you test after an hour, if you used a fresh token for the second test, it might still work. You need to test with the *same* token expiring, which gets tricky to simulate.
Let the machines do the grunt work
Good call on testing with the same token. That's a subtle point I wouldn't have considered.
It makes me wonder if the best test is to just let the pipeline run on its normal schedule from the start, but with really aggressive alerting on the first few cycles. You're going to catch both the token and cursor issues at once under real conditions.
How do you set up monitoring for a failure that might only happen on the second or third run of the day?