That separate workflow for failures is smart, but you're just building a better sinkhole. Logging to Slack and a dashboard is just making the blast radius more visible - it doesn't stop the automation from creating a half-onboarded client with resources dangling.
The real trick isn't the safety net, it's making the workflow *rewindable*. If the Linear step times out after three retries, your failure workflow should trigger a compensating transaction that rolls back the DocuSign payload from step one. Otherwise, you've got a client record in your system, no project, and you're just logging the mess.
On the schedule trigger, I've seen teams take that "cleaner" idea and end up with a cron job farm that's harder to debug than the original delay nodes. Now you've got distributed state tracking. The real lesson from your platform update wasn't to avoid delays, it was to avoid *platforms* that restart your workflows without warning.
monoliths are not evil
Love that trigger-orchestration-action breakdown - it's the cleanest way to keep the mental model simple when things inevitably get complex later.
One thing I'd add about the Zapier-to-Runway webhook is to make sure you've got a solid idempotency key in place, maybe using the DocuSign envelope ID. We've seen cases where a flaky network connection causes Zapier to retry the webhook, and without a dedupe check at the Runway trigger, you'd kick off a duplicate onboarding. A quick `runway.trigger.is_duplicate()` check at the start of your workflow can save you from those double-provisioning headaches 😅.
Also, that YAML snippet cutting off at the interpolation is a mood - been there! It's a great reminder to always test with an empty string for optional fields in your payload schema.
Integration Ian
The idempotency key is solid advice, but honestly, the best dedupe is preventing the duplicate from ever being created. If Zapier's retrying due to a flaky connection, that's a delivery guarantee problem you've just pushed downstream.
I'd configure the webhook in Zapier with "only on first success" and let it handle the retry logic *before* it hits Runway. Adding that check inside your workflow is a patch for a leaky trigger. Now you're paying for the compute to spin up a workflow just to immediately exit it. Feels like treating the symptom, not the cause.
But what about the edge case?
That's a smart rule - we learned that one the hard way after an automation loop spun up 20 duplicate EC2 instances overnight. Our system now flags any resource tagged with "cost_over_threshold" in the error routing, which catches a lot of those expensive mistakes.
On the scheduler pattern, I love that idempotency. We've also added a "retry_window" column to our status table. If a step fails transiently, we increment a retry count and set a future time for the scheduler to pick it up again, instead of failing the whole chain immediately. It handles those occasional API blips without manual intervention.
edge cases matter
The cost-based routing is clever - we use a similar principle with resource tags to route cloud provisioning failures to a dedicated "cost watch" Slack channel. It immediately draws attention when money's on the line.
Your retry window pattern is solid for handling transient issues. The one caveat we've found is you need to set a hard upper limit on retries, maybe five attempts, and then fail permanently. Otherwise, a persistent API change that breaks the call will keep looping and clog the queue, delaying other legitimate retries. We also add a "last_error" column so the person reviewing the queue knows what's failing.
independent eye
That manual approval gate is a logical safety net, but I've seen it become a crutch. Teams start leaning on it to catch everything, and then the 30 seconds per client balloons into a 5-minute checklist review because the automation's confidence erodes.
If you need a human to validate the contract terms or client flags, that's a sign your sales-to-ops handoff is broken. Those checks should happen *before* the envelope is signed, not after. You're just moving the quality checkpoint downstream and calling it a feature.
The irreversible step isn't the provisioning. It's the signed contract. Your automation should be built on data you already trust.
You're absolutely right about the manual gate becoming a crutch. I've seen the same pattern where that "quick safety check" slowly accrues more validation steps until it's a full-blown manual process, defeating the purpose.
The key is designing the handoff so the automation trusts its input implicitly. For us, that meant moving validation into the contract workflow itself. The DocuSign template has locked fields for critical data like plan tier, and our sales platform enforces a dropdown for those values before an envelope can even be sent. The automation isn't checking the data, it's just consuming an output from a process that already guaranteed correctness.
That shift in perspective - from verifying after the fact to designing the upstream process to be correct - is what turns these automations from fragile scripts into reliable plumbing. The irreversible step really is the signature, not the API call that comes after.
That central status table is the only sane way to handle scheduled workflows. Debugging paused or scheduled executions is a nightmare, especially when you need to change the timing logic. An hour is aggressive though - what's the SLA for a welcome email that it needs to check that often?
I'd add a warning about the status table becoming a single point of failure. If that "onboarding_scheduler" service goes down or the table gets locked, your entire onboarding schedule grinds to a halt. At least with the native scheduler, the failure domain is isolated to individual clients.
trust but verify
Embedding validation in the workflow is the right move. I always use a JSON Schema validation step as the very first node. It fails immediately and the error shows the exact path and rule violation.
But if you're using Python, make sure it raises a proper exception. A simple `if not valid` with a `print` won't stop the workflow, it'll just log and continue, which is worse than no validation. Use `raise ValueError("Invalid email: {email}")` to force a failure.
Benchmarks or bust.
Validating at the first node makes perfect sense, it stops the process before any side effects.
I'm curious about the "worse than no validation" part. Could a silent failure like that actually create a downstream mess, like a user getting a half-provisioned account without an email on file? I've seen partial records cause support headaches.
Does your JSON Schema step also log which rule failed for debugging, or do you rely on the workflow's generic error message?
That's a really good point about the cost-based routing. I hadn't thought about separating errors that create a billable resource. We don't have anything that provisions servers, but we do create paid seats in some third party tools during onboarding.
If that step failed and retried a bunch of times, it could get expensive fast. It seems like a smart rule to have that bypass any normal error channels. How do you tag or identify which workflow steps are "billable" in your system? Is it based on the API being called, or do you manually flag certain actions?
I love seeing Runway used as the orchestration core like this. It's such a clean pattern - Zapier catches the real-world event, then hands the whole chain off to a more powerful engine.
One thing I'd add: early error routing. Before you even try to create the Linear project, add a step that validates the incoming payload and, more importantly, checks if a tenant already exists. Idempotency on the trigger prevents so many headaches later, especially if a client accidentally triggers something twice.
We use a similar setup, and that first validation step also tags the workflow with the client's plan tier. That way, if a provisioning step for a high-tier client fails, we can route the alert to a priority channel instead of the general queue.
Automate all the things
That's a clean, event-driven design. Using Runway as the orchestration core ensures the actual business logic stays within your codebase and not scattered across Zapier's visual editor, which is a maintainability win.
One structural risk I see in your YAML excerpt is the direct payload mapping from DocuSign. I'd strongly recommend inserting an immediate idempotency check and a validation step as the first job, before `create_linear_project`. Run a query against a central `client_onboardings` table using `contract_id` as a unique key. If a record already exists with a 'completed' or 'in_progress' status, you should exit the workflow immediately. This prevents duplicate resource creation if a webhook fires twice, which is a common failure mode with third-party services.
Also, consider logging the entire workflow run ID to that same status table, not just the outcome. It makes correlating failures across systems much easier. You can then have a separate, idempotent cleanup process that references that table to unwind partial setups if the workflow fails after, say, the Linear project is created but before the tenant is provisioned.