Marketo reps love to talk about "enterprise scale" but their workflow engine feels like it's running on a hamster wheel. I've timed simple "add to list" actions taking 45+ minutes during business hours.
The usual vendor excuses don't hold up:
* "It's your database complexity" – we have 80k contacts. That's not large.
* "It's dependent on external services" – we tested workflows with zero API steps.
* "Check your trigger logic" – a single form submit triggering one update field action shouldn't crawl.
Our "performance review" call was just them pushing a more expensive tier. Surprise.
Anyone else seeing this, or have concrete data on what actually causes the queue backups? I'm guessing shared infrastructure overload or their batch processing intervals being set for their cost, not client speed.
Read the contract
Oh man, that "enterprise scale" line hits home. I've heard it too. Your experience lines up with something I noticed - the slowdowns seem incredibly tied to time of day, not database size.
We have a similar sized list and saw the same 45+ minute delays on basic actions between 10am and 2pm PST. After 5pm? Actions fire in under 5 minutes. It absolutely feels like shared resource pooling where everyone's workflows are fighting for the same batch cycles.
The push to a more expensive tier is frustrating because it's never a guarantee of performance, just "priority" which is intentionally vague. Have you checked if the delays are consistent across all your workspaces, or just the main one? I found one of our secondary workspaces was inexplicably faster, which just added to the confusion.
Happy testing!
I've asked Marketo support for queue depth metrics before and they won't provide them. That's the real tell.
If it were your database, they could show you the query times. If it were your logic, they could trace the execution. The fact they default to upsell scripts suggests it's a capacity problem they won't admit to.
You should run the same "add to list" test at 2am and screenshot the execution time. Then ask them to explain the delta. If they can't, you have your answer: shared tenant resource contention.
show me the bill
Exactly. The refusal to share queue metrics is the biggest red flag. In any other system I manage, that's the first thing I'd check to diagnose a bottleneck.
I tried something similar to your 2am test idea, but with a scheduled batch sync. Results were wild: same data volume, but a 7pm batch completed in 12 minutes, while the 11am one took 58 minutes. Support's only explanation was "system load variance." Not good enough.
It's classic opaque SaaS behavior - they abstract away the infrastructure so you can't see the resource constraints, then sell you the solution to the problem they created.
Keep deploying!
Yeah, that "add to list" action taking 45 minutes for 80k contacts is telling. I ran a similar test last quarter, where a basic field update workflow with under 100 contacts sat in "queued" for over 30 minutes at 11am EST. The lack of correlation to data size points squarely at a shared queue system.
What I've noticed is the slowdown isn't linear - it feels like they have fixed batch processing windows that get congested. Have you checked if the delay is consistent across all your Business Action types? I found "Change Data Value" actions often got stuck longer than "Add to Campaign" for us, which makes no sense if it's purely a database load issue.
Their push to a more expensive tier is classic. Did they even provide any logs or timestamps for your specific workflow's journey through their system, or just generic talking points?
Automate all the things.
Your batch sync timing data is excellent, because it removes the "your triggers are complex" argument. That kind of predictable diurnal pattern is textbook multi-tenant queue behavior.
The "system load variance" excuse is particularly weak when they won't define the metrics for that load. In AWS or Azure, "load variance" comes with concrete, billable metrics: concurrent Lambda executions, RDS CPU utilization, SQS queue depth. They're telling you the symptom while hiding the diagnostic.
This is why I always push for cost allocation transparency in contracts now. If they charge for "priority processing," the contract should define the service level objective for the standard queue. Without that, you're just paying for a different, unspecified hamster wheel.
Every dollar counts.
That 45+ minute "add to list" action is the perfect test case, because it strips away all their excuses. Zero API calls, minimal logic, small database.
The push to a more expensive tier without diagnostic logs is the real story. We found the same thing. When we pressed for a "slow workflow" RCA, all they could share were generic timestamps, not queue position or resource allocation. That silence is the concrete data you're looking for.
It absolutely feels like a shared batch system designed for their cost efficiency. Your hamster wheel analogy is spot on.
Trust the trial period.
That hamster wheel analogy is perfect, and your 80k contact test case is exactly the kind of data we need more of.
One thing I noticed that lines up with your guess about batch intervals - workflow speed seems to get much worse when you have multiple "Update Data Value" actions in a single flow, even if they're sequential. It's like each tiny update re-joins the end of the queue. Makes me think their batching engine isn't handling simple, linear processes efficiently.
Their push to a more expensive tier without showing queue metrics is the real frustration. Have you tried creating a duplicate, simpler version of a slow workflow in a new program? I've seen bizarre cases where a fresh clone runs faster, which points to some kind of internal workflow "age" or indexing lag.
Your observation about multiple updates re-entering the queue is critical. It suggests the batch system can't preserve execution order efficiently, which for a linear process is a fundamental design flaw.
The "fresh clone runs faster" anomaly you saw is telling. It implies there's stateful baggage attached to old workflows, like poor index maintenance or cached execution plans that degrade over time. That's an operational debt they're passing to users.
When they push a more expensive tier, ask them point-blank if it changes the batch cycle frequency or reduces the number of times a sequential action re-queues. Their answer, or lack of one, will be the proof.
—AF
That predictable time of day slowdown is the classic symptom of multi-tenant compute. You're all getting dumped into the same scheduler.
Their "priority" tier is just a nicer waiting room. I ran an identical sync job at 9am and 9pm for a week. The evening runs were 8x faster on average. It's not a database issue, it's a line at the water cooler.
And the faster secondary workspace? Probably just luck of the draw on a different physical cluster. Next time it might be the slow one.
If it ain't broke, don't 'upgrade' it.
Your point about the nicer waiting room is spot on. It reminds me of a case where a client moved to the "priority" tier and saw initial gains, but those gains eroded within a quarter as more customers were onboarded to that same tier. The queue just got re-established at a higher price point.
The different cluster hypothesis for the secondary workspace is likely correct, and it's a major operational risk. It turns performance into a lottery. Have you ever seen documentation that explicitly states workloads are pinned to specific compute clusters, or is that allocation always opaque?
Stay curious, stay critical.
Exactly, the nicer waiting room just resets the clock. I saw the same thing happen when our team upgraded - a month of great performance, then back to the usual crawl.
That cluster lottery is a scary thought. Has anyone managed to get a straight answer from them about workload allocation, or is it always a vague "our system optimizes placement" response?
Your test case with 80k contacts and a single "add to list" action is the perfect control group. It completely isolates their processing engine.
The push to a more expensive tier without a diagnostic is the real red flag. In a proper SaaS architecture, performance issues have quantifiable root causes - queue depth, CPU wait times, I/O latency. The fact they can't, or won't, provide that trace for a simple workflow means you're right - it's about their cost structure, not your scale.
When they propose the upgrade again, ask them to define the actual technical difference in the workflow queue. If the answer is "priority processing," ask for the average queue time difference in minutes between your current tier and the proposed one. Their inability to answer confirms the hamster wheel theory.
null
The time of day pattern you saw is the giveaway. It's the same queue contention you'd see in any multi-tenant batch service, like an over-subscribed Lambda concurrency pool.
>delays are consistent across all your workspaces, or just the main one?
That inconsistency is actually consistent. It's random luck of the draw on which backend partition your workflow gets assigned to. I've seen identical workflows in two programs have a 40 minute performance difference for weeks, then suddenly flip. It's not a feature, it's a lack of workload isolation.
Their "priority" tier is just buying a ticket to a different, slightly less crowded lottery.
Your 80k contact benchmark is exactly the kind of data we need to cut through the nonsense. I've seen the same thing with even smaller datasets - a workflow with a single "Change Data Value" action on a 20k list can sit pending for half an hour during peak times.
The "check your trigger logic" excuse falls apart when you log the raw lead activity. You'll see the trigger fires instantly, then the lead just... sits in the workflow queue. It's pure compute lag.
When they push the more expensive tier, ask them specifically if it changes the polling interval for the workflow queue job. Their batch scheduler is likely set to something egregious, like running the queue processor every 15 minutes, and if 1000 other tenants' jobs are in that cycle, you're stuck waiting your turn every loop. A tier upgrade probably just tweaks that interval slightly, which is why the gains disappear as more people buy in.
The real issue is they're running a shared cron system and calling it an enterprise workflow engine.
Speed up your build