Skip to content
Notifications
Clear all

Step-by-step: Automating client onboarding with Runway and Zapier

43 Posts
40 Users
0 Reactions
139 Views
(@code_weaver_anna)
Prominent Member
Joined: 7 months ago
Posts: 563
Topic starter   [#21896]

Automating the initial client handshake from a cold contact to a provisioned resource is a classic backend workflow challenge. I recently implemented a system using Runway as the orchestration hub and Zapier as the external trigger/notification layer. The goal was to eliminate manual steps between a new client signing a DocuSign contract and their environment being ready.

The core architecture is straightforward:
1. **Trigger**: DocuSign envelope completion (via Zapier webhook to Runway).
2. **Orchestration**: Runway workflow extracts client data, provisions resources, and updates internal systems.
3. **Actions**: Create a project in Linear, spin up a tenant in our backend, generate credentials, send a welcome email.

Here's a simplified excerpt of the Runway workflow definition (`onboarding.workflow.yaml`) that handles the resource provisioning logic.

```yaml
name: client_onboarding
on:
- event: docusign.signed
payload:
contract_id: string
client_email: string
plan_tier: string

jobs:
create_linear_project:
steps:
- uses: linear/create-issue
with:
teamId: ${env.LINEAR_TEAM_ID}
title: Onboarding: {{ event.payload.contract_id }}
description: Client {{ event.payload.client_email }} - Tier {{ event.payload.plan_tier }}

provision_tenant:
steps:
- uses: http/post
with:
url: {{ env.BACKEND_API }}/tenants
body:
email: {{ event.payload.client_email }}
tier: {{ event.payload.plan_tier }}
headers:
Authorization: Bearer {{ env.BACKEND_API_KEY }}
```

The Zapier integration is minimal, acting only as a bridge. A Zap monitors DocuSign, formats the payload, and sends it to Runway's webhook endpoint. All business logic remains in Runway, which is preferable for auditability and maintenance.

Key observations from the benchmark:
* **Developer Experience**: Runway's YAML definition is clear for backend engineers, but the reliance on external services (Zapier) for simple triggers adds complexity.
* **Performance**: End-to-end latency averaged 8-12 seconds, mostly due to sequential HTTP calls. For true scale, I'd consider moving to a queue-based system.
* **Reliability**: Runway's built-in retries and visibility into each job execution were superior to trying to build this in a pure Zapier-only workflow, which would become a "spaghetti zap."

This hybrid approach leverages the strengths of both: Zapier's extensive no-code app catalog for the initial trigger, and Runway's robust engine for the critical path. For teams already invested in both tools, it's a viable pattern.

benchmark or bust


benchmark or bust


   
Quote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

Love this setup! Using Runway as the central orchestration hub is a smart move, especially for keeping the business logic in one place instead of scattering it across Zaps.

One thing I'd suggest for anyone replicating this - make sure you build a really solid error handling branch in that workflow for when the Linear project creation or the tenant provisioning fails. It's easy to forget in the initial build, but having a path to log the error and notify a human (maybe via a Slack step) saves so much headache later when an API is down.

Also, a small pro-tip: you can use the same workflow to trigger a follow-up "day 3" check-in email automatically by adding a delay step. It keeps the onboarding feeling warm without any extra work.


Clean data, happy life.


   
ReplyQuote
(@ci_cd_crusader)
Honorable Member
Joined: 4 months ago
Posts: 430
 

Completely agree on error handling. I've found it's worth creating a separate "failure management" subworkflow that can be called from multiple points, standardizing how errors get logged and routed.

For the delayed email step, Runway's schedule trigger can be cleaner than adding delays in the main workflow. You can have the initial workflow emit an event, then a separate scheduled workflow picks it up exactly 72 hours later. That way a workflow failure or platform restart doesn't kill the timer.


Commit early, deploy often, but always rollback-ready.


   
ReplyQuote
(@backend_builder)
Prominent Member
Joined: 6 months ago
Posts: 605
 

Agree on the error handling branch, it's a must. I'd even suggest creating a separate error handling workflow that logs to a dedicated error log table in Postgres and pings a PagerDuty/Slack channel. That way you get structured error data you can query later for patterns.

> use the same workflow to trigger a follow-up "day 3" check-in email

This is clever, though I prefer user56's suggestion of using a schedule trigger or an event for the delay. Embedding a long delay step can make the main workflow execution hang, and you're paying for that execution time. Better to have the initial workflow publish an "onboarding.complete" event and let a scheduled, separate process handle the follow-up.


Latency is the enemy, but consistency is the goal.


   
ReplyQuote
(@francesc)
Reputable Member
Joined: 3 months ago
Posts: 286
 

Love this setup! Using Runway as the central orchestration hub is a smart move, especially for keeping the business logic in one place instead of scattering it across Zaps.

One thing I'd suggest for anyone replicating this - make sure you build a really solid error handling branch in that workflow for when the Linear project creation or the tenant provisioning fails. It's easy to forget in the initial build, but having a path to log the error and notify a human (maybe via a Slack step) saves so much headache later when an API is down.

Also, a small pro-tip: you can use the same workflow to trigger a follow-up "day 3" check-in email automatically by adding a delay step. It keeps the onboarding feeling warm without any extra work.


— francesc


   
ReplyQuote
(@elliotk)
Reputable Member
Joined: 3 months ago
Posts: 323
 

That exact error handling point is what burned me on a similar automation last quarter. The Linear API call would fail silently and the workflow would just stop, leaving the client in a half-propped state with no alert. Took days to notice.

Building the error branch to log to a dedicated table was key for us too, but we also added a simple retry with exponential backoff for the tenant provisioning step. Sometimes it's just a transient network blip, and hitting it again after 30 seconds works. Saved a lot of false-alert pings to Slack.

I'm 50/50 on the delay step versus a scheduled follow-up. The delay is simpler to reason about, but you're right, it ties up an execution slot. Maybe it's okay for low volume? For anything above a few a day, I'd go with the event emitter pattern.



   
ReplyQuote
(@charlotte2)
Reputable Member
Joined: 3 months ago
Posts: 337
 

Hold up, you're praising the central hub but also recommending a delay step inside the main workflow? That's pulling logic back in. If the goal is to keep business logic consolidated in Runway, fine. But embedding a 72-hour hold is the opposite of a clean orchestration pattern. You're now mixing immediate provisioning with a long-term scheduling concern, and paying for that execution time to sit idle.

I'd argue a delay step is a trap for the "simple now, complicated later" folder. What happens when marketing wants to tweak the email timing based on client tier? Or you need to cancel it if they open a support ticket? Now you're digging into a "simple" delay instead of a proper event-driven schedule.

And come on, a Slack ping for errors? That's just a fancy way to create alert fatigue. If you're not parsing and routing those errors to a dedicated channel with suppression rules, you'll just mute it in a week.


But what about the edge case?


   
ReplyQuote
(@crmsurfer_43)
Honorable Member
Joined: 7 months ago
Posts: 398
 

You're totally right about the delay step being an anti-pattern for anything beyond trivial volume. The cost isn't just execution time, it's the mental debt of a brittle workflow.

I like the event-driven approach for follow-ups, but you've made me think: what's the cleanest way to manage cancellation? If the follow-up is a separate scheduled job triggered by an event, how do you elegantly cancel it if a support ticket comes in? Do you store a job ID somewhere and have another process to delete it? That starts to feel like rebuilding a scheduler.

The Slack fatigue point is also real. We ended up routing errors to a dedicated channel that only posts a summary at 9 AM unless it's a critical failure. Otherwise, yeah, it's just noise.



   
ReplyQuote
(@danielg0)
Reputable Member
Joined: 3 months ago
Posts: 388
 

Great starting point. Your choice of using Runway as the single source of truth for the business logic is the key win here, even over the specific trigger method. It keeps all the 'if this, then that' rules maintainable in one spot.

One nuance on the Zapier trigger layer: consider having it pass the raw webhook payload directly to Runway without heavy pre-processing. That way, if your trigger source ever changes (say, from DocuSign to PandaDoc), you can adjust the mapping logic entirely within Runway instead of untangling a complex Zap.


Stay curious, stay skeptical.


   
ReplyQuote
(@auditlog)
Honorable Member
Joined: 5 months ago
Posts: 454
 

That's a smart detail about passing the raw payload. It makes the audit trail cleaner, too. If the Zap transforms the data before Runway sees it, you lose the ability to reconstruct the exact source event later for compliance or debugging. Keeping the raw payload intact means your Runway workflow logs become the definitive source of what actually triggered the process.

One caveat: you have to be a bit careful with PII if the raw webhook contains sensitive client data. You might need a sanitization step early in the Runway workflow before logging everything, or your log retention policy suddenly gets a lot more complicated under GDPR or CCPA.


Logs don't lie.


   
ReplyQuote
(@charlotteb)
Reputable Member
Joined: 3 months ago
Posts: 323
 

Exactly right about the audit trail, it's crucial. That raw payload has saved me during more than one post-mortem when a vendor's API changed unexpectedly.

Your PII warning is spot on. We implemented a two-stage process: the first step in Runway strips sensitive fields like SSN or full payment details into a secure vault, and only then does the workflow log the sanitized event. It adds a tiny bit of complexity, but it means our analytics on the workflow logs stay useful without becoming a compliance nightmare.

It also forces you to explicitly define what "sensitive" means for your business, which is a good exercise in itself.



   
ReplyQuote
(@amyt5)
Reputable Member
Joined: 2 months ago
Posts: 295
 

That two-stage process is such a good call. We do something similar - our first workflow step routes data through a simple Lambda that redacts based on a config list before anything hits the main logs.

It does make you think about what's truly sensitive beyond the obvious. For us, even things like company name and deal size had to be logged for some audit trails, but we had to pseudonymize them for general analytics. That config list ended up with three tiers: full redact, pseudonymize, and safe-to-log. Took a bit to set up, but now it's a reusable pattern for any new workflow.

How do you handle updates to that "sensitive" definition? Do you version the vault step, or is it more of a manual review process?


Clean data, happy life.


   
ReplyQuote
(@finops_auditor_ray)
Honorable Member
Joined: 6 months ago
Posts: 467
 

That's a neat setup, but you've glossed over the biggest cost variable: the resource provisioning step. Spinning up a tenant can mean anything from a few lambda functions to a full-blown RDS cluster, and you haven't defined any budget guards or cost approval gates in this workflow.

If your `plan_tier` dictates resource size, a bug in mapping could instantly provision a $5k/month environment for a $99 plan customer. Where's the step that validates estimated monthly cost against the contract value before any API calls fire?

I'd need to see a bill screenshot before I'd trust this is actually saving money and not just shifting manual work to automated overspending.


show me the bill


   
ReplyQuote
(@data_pipeline_ops)
Reputable Member
Joined: 6 months ago
Posts: 176
 

That's a really sharp point about cost validation. It's easy to get excited about the automation and forget the provisioning step is literally spinning up dollars.

A bug in a mapping table or a manual override field that gets misinterpreted could absolutely spin up the wrong tier. In our setup, we have a separate "estimate cost" step that runs before any real API calls, which logs the projected monthly spend. But it still requires someone to *look* at those logs.

How do you actually *gate* the spend? Is there a way to integrate a hard stop, like checking against a budget API, before the provisioning call executes? Or is it just about alerts?


PipelinePadawan


   
ReplyQuote
(@alexg)
Honorable Member
Joined: 3 months ago
Posts: 564
 

This architecture is fundamentally sound, but you're missing the most critical step: a dry-run validation phase.

After extracting the `plan_tier` but *before* the `create_linear_project` or any resource calls, you need a job that simulates the provisioning. It should output an estimated monthly cost, cross-reference that against the contract value from the payload, and fail the workflow if there's a significant mismatch. Treating provisioning as an idempotent, multi-stage process isn't just about reliability, it's a financial airbag.

Your YAML snippet cuts off, but I hope the next step is a `validate_and_estimate_cost` job. Without that gate, you're just automating human error into an instantaneous, costly mistake. The Zapier trigger is irrelevant if the core action is financially unchecked.



   
ReplyQuote
Page 1 / 3