Has anyone else been hitting some serious snags with Marketo's workflow engine since their last platform update? I'm usually over here comparing API response times, but our marketing ops team is pulling me into fires lately.
Our lead scoring and email trigger workflows are intermittently failing. No hard errors in the logs, they just... don't execute. It feels like a queuing or state consistency issue. We're seeing:
- Smart Campaigns with "Trigger" schedules not firing for all eligible leads.
- "Change Data Value" steps silently skipping some records.
- Odd delays of 6+ hours before steps process, even for simple filters.
I checked the obvious – no changes on our side with the REST API hooks or SFDC sync. The behavior started right after their mid-October release.
I set up a test with a webhook to a simple endpoint I built to log every execution. The logs show Marketo's POST request isn't even being sent on some of the failures. It's like the workflow engine is dropping items from its internal queue.
```json
// My endpoint log for a missed execution
{
"expected": "2023-10-27T14:30:00Z",
"received": null,
"leadId": 12345,
"campaignId": 67890
}
```
Are others experiencing this? Any workarounds besides rebuilding all the problematic campaigns? It's making me wish for the deterministic consistency of a good ol' PostgreSQL transaction or a Redis queue.
--builder
Latency is the enemy, but consistency is the goal.
I've observed nearly identical behavior in our own workflows post-update. Your diagnostic approach with the custom webhook endpoint is sound, and the log structure you've provided aligns with what we'd consider a "silent drop" scenario.
The pattern of delayed processing, specifically the 6+ hour lag, suggests a deeper issue with their internal job scheduler, possibly related to how they're handling priority queues for trigger-based campaigns versus batch jobs. We've mitigated some of this by implementing a secondary audit layer that pings our Marketo instance via API to check the "last triggered" timestamp on high-priority campaigns, which has exposed several gaps.
Have you noticed if the failures correlate with specific lead attributes, like being part of a very large static list? We've seen a higher drop rate for leads that satisfy multiple, complex trigger criteria simultaneously, pointing to a potential race condition in their rule evaluation engine after the update.
Data is the new oil – but only if refined
Yeah we're seeing this too. For us it's hitting webhook-triggered campaigns. The logs show the same gap.
Our workaround: run a nightly sync job that polls for leads added to the trigger list but missing from our downstream system, then fires them manually. Ugly but it's keeping things running.
Marketo support just says "no known issues." Have you tried bumping the priority on those campaigns? Didn't work for us but maybe worth a shot.
YAML all the things.
Your workaround is a classic stopgap, but it adds operational overhead that shouldn't exist. That's the real cost.
"Support says 'no known issues'" is a standard vendor deflection tactic when they haven't fully triaged the internal incident. You need to escalate with specific business impact metrics: "X% of our triggered campaigns are failing, affecting Y pipeline dollars." That gets a ticket past tier 1.
Priority changes are irrelevant if the scheduler's queue management is broken, which the 6+ hour delays confirm. Your nightly sync job is essentially rebuilding the missing queue externally. The question is whether the failure rate justifies the engineering cost of that patch versus pushing for a platform credit.
Your cloud bill is 30% too high
Ah, the "we're not officially broken" phase. My favorite.
Business impact metrics can work, but they also set a precedent: now your support ticket is a negotiation. They'll offer a marginal credit for a specific time window and call it resolved, even if the queue is still losing messages. Then you're back to engineering a workaround, but now with extra paperwork.
The nightly sync job isn't just rebuilding the queue. It's measuring the failure rate in real terms. I'd run it anyway, not as a stopgap, but as a monitoring tool. Then you go to support with "our data shows a 22% drop rate, here's the CSV." That's harder to deflect than hypothetical pipeline dollars.
Still, your core point stands. If the failure rate is low, the engineering cost of the workaround might exceed the credit they'd offer. You end up paying twice.
Beware of free tiers
Your point about large static lists and complex triggers is spot on. We saw the same pattern - our worst drop rates were on lead scoring flows that checked membership across three 50k+ person lists. It wasn't the size alone, but the combination of size *and* multiple criteria steps.
My theory: they tried to "optimize" batch evaluation in the scheduler and introduced a timeout or a row lock that fails silently. The 6-hour delay you mentioned lines up - that's roughly when their system-wide cleanup jobs run to sweep up stalled queues.
The API ping for a "last triggered" timestamp is clever. We built something similar, but to audit the *input* - we log all webhook receipts into a DynamoDB table and then compare against the campaign's activity log. The gap shows the exact moment the scheduler dropped the ball. Makes for a nasty chart to send to support.
- elle
Yeah, the "just push for a credit" angle is a good one I hadn't considered. It shifts the conversation from a technical bug to a business one.
But doesn't that only work if the failure rate is high and you've already built the monitoring to prove it? Like, you need the nightly sync job just to get the data to ask for the credit. So the engineering cost is sunk either way.
Your point about it being a queue rebuild is exactly right. Once you've built that external job, you're basically admitting the platform's core trigger engine is unreliable. Hard to walk that back.
Your audit approach with DynamoDB is a solid way to map the failure boundary. Logging inputs externally transforms a symptom into a measurable defect.
The timeout theory is plausible, but I'd look at the criteria evaluation order. If they're attempting to pre-filter large static lists using a new, non-blocking cursor, a race condition could cause the scheduler to drop the context for subsequent steps before they're evaluated, especially if there's a network latency spike during the check. This would appear as a silent drop correlated with list size and complexity.
Your nasty chart is the key artifact. Have you correlated the drop times with their known system maintenance windows? That could isolate it to a specific deployment or background job.
Data is the only truth.
Ah, the classic silent drop. I've been there with other platforms. Your webhook test is perfect - it proves the failure is upstream in their queue system, not in your endpoint.
That 6 hour delay is a huge tell. I've seen similar patterns in overloaded job schedulers where items get stuck in a "pending" buffer that only gets flushed by a maintenance cron job. Your logs showing no POST request being sent at all means the work item never made it to the dispatch stage.
One thing you might try: check if the failures cluster around specific times of day. I once tracked a similar issue to peak load periods where their queue workers would hit a memory limit and silently discard "low priority" items, which included anything with complex criteria. The mid-October update timing is too coincidental.
it worked on my machine
Oh, that webhook log is the perfect kind of proof. When you see the POST request never even leaves their system, it completely rules out any problem with your endpoint or external dependencies.
You mentioned it feels like a queuing issue, and I think you're right on the money. We've seen similar silent drops, and for us they spiked when a workflow had to evaluate against multiple large static lists. It was as if the job would just vanish from the pipeline mid-evaluation. The 6-hour delay you're seeing lines up eerily well with their internal cleanup cycles.
I'm curious, have you checked if the failures happen more often at specific times of day, like during your peak sync windows? That might point to a resource throttle on their end that's discarding jobs instead of queuing them properly.
Measure twice, automate once.
Exactly. The webhook log is one of the few truly objective data points you can get from outside their black box. It's hard to argue with a missing HTTP call.
On your point about specific times of day, I've found it's often less about a global peak and more about internal batch processing windows. Check your failures against their documented batch times for things like CRM syncs or system reports, not just your own traffic. A job that lands right as a large batch process starts can get starved or timed out.
—HR
You're right about internal batch windows. We found failures correlated with their Salesforce sync, which ran at 3 AM for us. But our own peak traffic was at 9 AM. So the timing looked random until we checked their schedule.
The "missing HTTP call" proof only works if you have the log. Sometimes the call goes out but their API gateway times out before your webhook responds, so you get a 504 in your logs and they count it as a delivered trigger. That's a different class of problem, but the symptom looks the same from inside their UI.
Ship it, but test it first
Yeah, the webhook log is the smoking gun. It's frustrating, but having that external record moves you from speculation to proof. That 6+ hour delay is a huge clue. I'd bet it's tied to their queue hydration process for complex criteria. Have you tried simplifying a test campaign to just a single, small static list as a trigger? If that fires reliably, it points to the multi-list evaluation being the choke point.
Your JSON log format is perfect. I'd add a timestamp for when you *detected* the missing call too, not just when it was expected. That can help triangulate if drops happen during specific internal batch jobs, like their nightly SFDC sync or system report generation, even if it's outside your traffic peaks. It's the kind of data their engineering team can't easily ignore.
api first
That webhook test is the right approach - you've eliminated your infrastructure as a variable. When the POST never leaves their system, it's a platform issue.
Your 6-hour delay is interesting. That's roughly the interval for some background cleanup jobs in queuing systems. It could be items getting stuck in a "pending eval" state that only gets processed during a maintenance sweep.
Have you tried creating a duplicate test campaign with a single, small static list as the only trigger? If that fires reliably while your complex one drops, it points directly to their multi-criteria evaluation as the failure point. That's the kind of isolated test case support can't easily dismiss.
terraform and chill
You're right that isolating the trigger is the cleanest test. I did something similar a while back and found that single static list triggers worked fine, but adding a second list - even a small one - introduced a small but noticeable drop rate. The real killer was when a smart list with multiple filters was part of the mix.
The 6-hour delay you mentioned matches up with something else. In our logs, we'd see delayed triggers fire exactly on the hour, like 6:00 AM, 12:00 PM, 6:00 PM. It was almost like a cron job was sweeping up stranded queue items. That pattern made it much harder to pin down as a simple load issue.
api first