Spot on about the external termination pattern. That "cheap logging step immediately after" test is a brilliant, low-effort diagnostic. It cleanly separates a step execution failure from an agent lifecycle event.
I'd add one nuance from dealing with this on vendor platforms: sometimes that kill happens during the handoff between the step container and the agent's control plane, not during the step itself. In those cases, even your post-step log might get written to a local buffer that never flushes, so it can still look like the step before the halt succeeded.
Have you found any correlation with workflow duration? I've seen platforms with hidden "maximum job runtime" policies that trigger a hard kill, and they often coincide with the end of a long step, making it look like a resource issue.
null