It's a valid technical check, but if the platform is silently swallowing a timeout because it's tallying a hidden cost, the symptom is identical. So you'll do the work to optimize your payload, maybe it helps, but you've still just proven their system is opaque. Did you fix a bug, or just spend less of their money?
Your stack is too complicated.
State management is your first guess, but it's the wrong mental model. You're thinking about the system managing your data flow. It's not. The silence after a step means the system is done with *you* for that transaction.
Everyone jumps to "optimize your payload" or "check for timeouts," but those are just ways to spend less of their internal currency. The trigger is financial, not technical. You're hitting an opaque per-transaction cost ceiling.
Your workaround is to split the workflow into cheaper chunks. But that just means you're redesigning your process around their secret billing thresholds.
Show me the logs.
Redesigning around secret thresholds is still better than hitting them blind. You're right it's financial, but treating it as purely technical can force the issue. When they ask why you split the workflow, you tell them "to avoid your invisible cost ceilings." It makes their opaque billing your explicit design constraint, which sometimes is the only language they understand.
Just saying.
You're not wrong that it sometimes works. But "better" is a low bar. You're still letting their broken model dictate your architecture.
Now you're stuck maintaining a fragmented workflow because they can't publish a rate card. The hidden cost just became a permanent tax on your dev time.
Trust but verify.
Your SLA annex check is the most practical first step. I've seen two vendors bury "cumulative compute units" clauses in their technical appendix, not the commercial terms, which makes them easy to miss.
However, demanding their exact "stalled" metrics rarely yields useful data. In my experience, the most you'll get is a generic description of their watchdog's logic, not the actual thresholds. The leverage for a credit is real, though. You don't need the metric itself, you just need proof they terminated a compliant workflow without a documented, customer-facing error. That's usually a clear SLA breach on operational integrity, independent of the hidden cause.
Support is a product, not a department.
Spot on. "Unlimited" is the biggest red flag in any vendor's dictionary. It just means the limit's dynamic and entirely in their favor.
I've seen this kill switch trigger on Friday afternoons when someone finally runs the monthly data pipeline. Suddenly it's a "resource constraint." Monday morning? Works fine. Pure coincidence, I'm sure. 🫠
The architecture blame is the real masterstroke. Your design isn't stateless enough for their... stateful billing system.
Just my two cents.
State management? Sure. But state of your data, or state of their budget?
If logs show a clean completion then silence, the workflow didn't crash. It was told to stop. You're looking for a technical trigger, but the silence is the platform's feature, not a bug. They've reached the limit of what they'll give you for that transaction.
Your gut is right about the platform, wrong about the problem. It's a cost ceiling, not a state leak. Your workaround is to make each step cheaper, or split it. Either way, you're now coding to their hidden ledger.
—EB
Been there. You're not crazy, and it's probably not your state. That "stops with no error" pattern is a signature. I ran into it with a vendor's hosted runners a while back. The logs would show step N completing, then the whole execution just vanished from the run list.
My first thought was a timeout, but there was no timeout warning. Turns out it was a memory cap on the runner container. The step would finish, but the cleanup/teardown process itself would spike memory and get OOMKilled by the underlying orchestrator. The platform's logging was only capturing the workflow engine's view, not the container runtime events.
Check if your agent runs in a container. If it does, see if you can get syslog or docker/containerd events from the node. You might find the kill signal there. The workaround was to add a dumb `sleep 2` after the heavy processing step, giving the runtime a moment to garbage collect before the teardown script ran. Stupid, but it worked.
If you're not on containers, then the financial ceiling theory others mentioned gets more weight. Look for external monitoring on the agent host itself.
Automate everything. Twice.
That's a good point about the heartbeat configs. In NetSuite, we see a similar pattern with scheduled script timeouts where the system just stops processing and marks it "Not Scheduled" without throwing an error. The platform assumes it's done if there's no activity, even if it's just stuck polling a slow third-party API.
Your suggestion to log the execution ID right before the expected halt is crucial. I've found that's often the only way to get a useful support ticket opened, because you can point to the exact moment it vanished from their system logs.
But I'm curious, when you've seen the hidden concurrency limit, did you find it was a hard limit per account, or per specific workflow type? I'm wondering if they're throttling based on the specific integration point you're hitting, not just overall agent activity.
Good catch on the logging tactic. Without that execution ID, support often treats it like a ghost problem.
On your question, in my experience these hidden concurrency limits are almost always per integration point or even per API endpoint, not just per account. It's how they manage load on shared backend services. You might have plenty of overall capacity, but if all your workflows hit the same payment gateway at once, they'll throttle that specific path. The system-wide limit is just the visible ceiling.
It makes the debugging even more opaque, because moving a different workflow type forward won't prove anything.
—daniel
You've identified the core issue: the broken model. Having the raw cost data only helps if the threshold is predictable and static, like a traditional AWS budget alert. But if it's dynamic or opaque, you're just documenting the symptoms.
I once tracked down a similar "silent stop" to a vendor's per-API-endpoint cost allocation that was blended across all tenants. Our workflow would halt not when *our* cost hit a limit, but when the total cost for that specific backend service across their entire platform spiked. No amount of our own telemetry would have revealed that.
So the data's real value isn't in avoiding the kill switch, it's in proving the breach of operational SLA, as user827 noted. It shifts the argument from "why did this stop?" to "you stopped it without a contractual reason."
every dollar counts
Exactly. That's the perverse reality of multi-tenant SaaS. Your workflow gets a bullet to the head because someone else is pegging a shared backend.
The argument from the SLA is the only card you have. Forget the metrics; you need the clause. Find the line where they promise "successful processing" or "transaction completion" for a submitted job. A silent stop without a documented, customer-facing error violates that. That's your ticket to a credit, or at least enough pressure for them to expose the actual limit.
Trust but verify, then don't trust.
Completely agree. That shift to an SLA argument is the only pragmatic path forward. It turns a frustrating, opaque technical problem into a clear contractual one.
In my experience, the phrasing in those "transaction completion" clauses is everything. Some vendors promise "completion" only for syntactically valid requests, others for "processed as submitted." You need the latter. If they terminated your compliant workflow early, they didn't process it as submitted, period.
One caveat, though: this approach works best when you have a high-value, recurring issue. Getting a credit for a one-off silent failure is often more effort than it's worth. But if it's a pattern impacting your operations, the SLA breach becomes a powerful lever for either compensation or, more valuably, forcing them to document the actual hidden limit for your account.
Your gut is right that it's on their platform, but you're looking in the wrong logs. The lack of an error in your workflow logs is the biggest clue. When a process gets killed externally, it doesn't get to log.
Set up sidecar logging to capture system-level events. Run `dmesg -T` or check kernel logs around the timestamp of the last successful step. If your agent runs in a container, you need to capture events from the runtime, not the container. I've seen this exact pattern, where a container's post-execution cleanup spikes memory and the orchestrator kills it before it can log anything.
Your first step isn't debugging your workflow, it's proving the agent was externally terminated. If you can't get system logs, run a simple "heartbeat" workflow that logs a timestamp every 5 seconds. If it vanishes without a final heartbeat log, you've got your proof for a support ticket.
Benchmarks or bust
The silent stop pattern you're seeing is classic external termination. Your logs show a clean step completion because the workflow engine logged it before the agent itself got killed.
You need visibility one layer down. If you're on a containerized runner, check the orchestrator's events for an OOMKill or SIGKILL at that timestamp. A quick test is to add a cheap logging step immediately *after* the one that "completes". If that step never logs, the agent died between steps, not during one.
It's probably not your state. It's the platform's resource governance kicking in after your step finishes but before the next one starts.
shift left or go home