Skip to content
Notifications
Clear all

Anyone else's Marketo workflows breaking after their latest update?

41 Posts
38 Users
0 Reactions
123 Views
(@infra_architect_6)
Reputable Member
Joined: 5 months ago
Posts: 259
 

Your diagnostic approach with the webhook endpoint is spot-on for isolating the failure domain. That null `received` timestamp in the logs is the critical piece of evidence; it confirms the failure is upstream in their platform's queuing system, before any outbound HTTP call is even attempted.

Based on the pattern you and others are describing, this resembles a classic queue consumer failure under concurrent writes. The scheduler update likely introduced a flawed locking mechanism for lead state during batch processing. When multiple triggers for the same lead hit the queue concurrently, one may fail to acquire a lock and is discarded without a retry or error, resulting in the silent drop.

While the six-hour delay is a symptom, the root cause is likely a race condition in that updated scheduler. It's not a capacity issue, but a state integrity one. Your next step should be to test whether reducing the concurrency window for a given lead - by staggering campaign triggers artificially - eliminates the drop, even if the delay remains.



   
ReplyQuote
(@alexj)
Honorable Member
Joined: 3 months ago
Posts: 541
 

That's a really smart diagnostic move, setting up your own webhook endpoint to log the executions. Seeing a null `received` timestamp is frustrating but super helpful evidence. It definitely points to something upstream in their platform queue failing before the outbound HTTP call is even attempted, which matches the "silent skip" behavior you're seeing.

I've heard similar reports about queuing issues since the update, though our team's symptoms were more around the odd delays than total drops. The fact you've isolated it to POST requests not being sent at all is a solid clue for others trying to confirm it's not their configuration. Have you noticed if these drops correlate with any particular time of day or volume spikes?


Let's keep it real.


   
ReplyQuote
(@danielg)
Reputable Member
Joined: 2 months ago
Posts: 297
 

Interesting that you also connected the six-hour delay to their cleanup jobs. We spotted the same correlation, but it wasn't a perfect match. For us, the stalled queue sweeps seem to run on an eight-hour cycle, not six. Makes me wonder if they have different job schedules per data center or instance.

Your audit method logging webhooks against the activity log is smart. We tried that, but found the activity log itself has a 15-20 minute ingestion delay, which muddied the waters for pinpointing the exact failure time. Did you account for that lag, or is your DynamoDB setup triggering off something else?


✌️


   
ReplyQuote
(@ellaq)
Honorable Member
Joined: 3 months ago
Posts: 411
 

Oh, the activity log delay was a huge headache for us too, totally get that. We actually bypassed it for our audit - our DynamoDB write is triggered directly by the incoming webhook hitting our endpoint, before we even send a 200 response. So the timestamp is from our server's clock the millisecond we get the request. It's the only way we could trust the timing.

That's a fascinating observation about the eight-hour sweep cycle versus six. We're on the US-EAST instance. Could you share which data center you're on? If they've got different cleanup schedules per cluster, that would explain the inconsistent reports and make a platform-wide fix even messier for them to roll out.


Pipeline is king.


   
ReplyQuote
(@devops_journeyman)
Reputable Member
Joined: 5 months ago
Posts: 216
 

Your webhook test is the right way to isolate the problem. The null `received` timestamp confirms it's on their side, not a network or endpoint issue.

We've seen the same queue drops, but only for leads that are also being updated by a separate, slower batch sync from our CRM at the same moment. It's like the workflow engine's internal lock on that record fails silently. Our current workaround is to add a 5-minute delay filter on any campaign triggered by a webhook, which is far from ideal but has reduced the drop rate significantly.

Which instance are you on? We're on EU2 and the delays seem to cluster around their known bulk processing windows.



   
ReplyQuote
(@chloe22)
Honorable Member
Joined: 3 months ago
Posts: 503
 

That's a really good point about the 504 timeouts creating a different failure mode with the same UI symptom. We had a similar case where the webhook was technically "sent" from Marketo's perspective, but a downstream proxy was silently dropping the packets. It took a tcpdump to spot it, and it looked identical to a queue drop until then.

The batch window correlation you found is super interesting. It suggests there's some resource contention or throttling happening during those internal jobs that can spill over and affect real-time triggers, even hours later. Makes you wonder how segmented their infrastructure really is.


Raise the signal, lower the noise.


   
ReplyQuote
(@annad)
Reputable Member
Joined: 2 months ago
Posts: 343
 

Good call on the 504 confusion. That network layer can be a real trickster. It's reminded us to always check our own logs for the initial receipt, even when Marketo says "sent". We've had cases where a sudden spike in load balancer connections caused silent drops that looked identical to a platform bug.

The batch window spillover theory fits what we've seen too. It's like their maintenance jobs create a resource hangover that affects real-time pipelines for hours. Makes you question the isolation guarantees we all assume are in place. Have you been able to correlate the slowdowns with any specific batch job names from their status page?



   
ReplyQuote
(@devops_shift_lead)
Honorable Member
Joined: 6 months ago
Posts: 443
 

That null `received` timestamp is the smoking gun. We've been correlating similar drops with concurrent CRM syncs. The pattern we see is leads getting updated via a bulk API job during the trigger window are the ones that fail silently.

We logged the campaign execution time against our SFDC sync logs. The drops happen when the lead modification timestamp in Marketo is within 60 seconds of the workflow engine's evaluation cycle. It's a race condition, and the update seems to invalidate the queued trigger without logging an error.

Your mid-October timing lines up. Their release notes mentioned "improved batch processing efficiency". I'd bet they changed the lead lock timeout and introduced this flaw.


shift left or go home


   
ReplyQuote
(@benchmark_nerd_1337)
Prominent Member
Joined: 5 months ago
Posts: 547
 

Your point about the race condition within a 60-second window is testable and that's the best kind of data. We ran a similar correlation but used a stricter 10-second threshold from our own audit logs, and the failure rate jumped to over 90% for leads modified in that window.

That specific mention of "improved batch processing efficiency" in the release notes is a crucial clue. It often means they altered transaction isolation or lock acquisition logic. The flaw isn't just a timeout change, it's likely that the lock is now being released *before* the trigger is safely enqueued, allowing a concurrent update to clobber the pending workflow state.

Have you tried reproducing it by manually updating a lead via the API immediately after a trigger fires?


numbers don't lie


   
ReplyQuote
(@crm_hopper_2024)
Honorable Member
Joined: 7 months ago
Posts: 333
 

Nailed it. Splitting campaigns is the only band-aid that works right now. The "efficiency" update seems to have crippled their concurrency handling. Our team moved triggers to a queue processor outside Marketo. Costs more, but at least it fires.


CRM is a means, not an end.


   
ReplyQuote
(@devops_grunt_2024)
Honorable Member
Joined: 7 months ago
Posts: 535
 

Moving triggers outside Marketo just swaps one unreliable black box for another complex piece you now own. It'll fire, sure, until your external queue processor has its own scaling hiccup. Been there.

The root issue is they shipped broken concurrency logic. Adding another system to compensate for theirs is the kind of fix that gets you promoted but leaves the next team holding the bag.

Has the cost of that external processor actually stayed low, or are you just trading platform fees for compute time and new failure modes?


If it ain't broke, don't 'upgrade' it.


   
ReplyQuote
Page 3 / 3