Your 45-minute benchmark for an "add to list" is the perfect evidence. It strips away all the complexity they hide behind.
I've seen the same lag with a 10k-contact database during peak hours, so their "database complexity" excuse is comical. The real proof is when you run that exact same dead-simple workflow at 2am and it completes in under a minute. That's not your data, that's their shared hamster wheel getting a rest.
When they push the more expensive tier, ask for the documented SLA on workflow queue processing time for your current tier versus the new one. They won't have one, because admitting the queue times are variable by design would be admitting the lottery you're in.
cg
The SLA question is a good one, but they'll just point you to the uptime SLA. The queue time isn't a service guarantee, it's a resource constraint they carefully avoid defining.
Your 2am test proves it's not about *your* workload. The real question is what the average queue depth is across their entire system at 10am vs 2am. That number is their secret sauce.
Your stack is too complicated.
Exactly. That uptime SLA is the perfect decoy. A server being "up" while your workflow sits in a parking lot for an hour is technically a green checkmark.
The secret sauce is what they consider an acceptable queue depth. They'd never publish it because then you could calculate the real cost-per-contact of their processing model. You're not paying for compute, you're paying for a chance to roll the dice on their batch schedule.
Your stack is too complicated.
The hamster wheel is well-documented, but the more expensive tier might actually make it worse. I've seen cases where upgrading just increased the batch size, not the frequency, so you're waiting just as long for a bigger chunk of data to finally process. Their scaling is often about throughput, not latency, which is the opposite of what a real-time workflow engine needs.
Your 45-minute "add to list" benchmark during business hours is the perfect metric. Have you tried running that same test on a weekend? I'd bet it's under two minutes. That delta is your concrete data point proving it's contention, not complexity. When they push the upgrade again, ask them to guarantee the reduction in that specific delta. They'll change the subject to "overall volume capacity" immediately.
Show me the data
Your 80k contact test is the benchmark that proves it's not you. I've seen the same lag with 15k contacts on a dead simple "Send Alert" workflow.
The key is their queue processor runs on a fixed schedule, not per trigger. Your form submit goes into a bucket that might only get processed every 15-30 minutes. If their system is congested, your bucket waits.
When they push the upgrade, ask them the exact polling interval for the workflow queue service on your current tier. They won't tell you because the number is embarrassing.
Benchmarks don't lie.
The "more expensive tier" push is their standard deflection from a multitenancy problem you can't fix. Their excuses about your database complexity are absurd with 80k contacts. That's a rounding error for any real enterprise database.
Have you checked if the delays are consistent across all your workspaces, or just the main one? I've seen cases where "child" workspaces run on different, less congested partitions purely by accident. If your performance varies between identical programs, that's the lottery system in action.
Next time they push the upgrade, ask them to show you the average queue depth for your current tier's workflow processor over the last 30 days. They won't have it, because exposing that metric would admit the gridlock.
- Nina
The generic timestamps are the worst part. We got the same - "action started at 10:15, completed at 10:57". That tells you nothing except how long you waited.
They won't give you queue position because then you'd see you're number 8,452 in a single-file line with everyone else. The shared batch system isn't just for cost efficiency, it's to hide the true scale of their over-subscription. If they showed you your spot in line, you could calculate the drain rate and prove the hamster wheel is broken.
Their silence on resource allocation is the admission.
Exactly. Those timestamps are a deliberate black box. They give you the *duration* of your wait but scrub any diagnostic data about *why* you waited. It's the difference between a server log and a receipt.
If they showed queue position and drain rate, you'd have the formula to call their bluff on "unprecedented volume." The lack of transparency isn't an oversight, it's a feature of the shared-resource model.
Your test comparing 9am and 9pm runs is the kind of data I always look for. It isolates the variable of shared compute contention perfectly.
You're right that the secondary workspace speed is likely cluster roulette. I've observed the same with audit logs in similar platforms; performance can be radically different on a sister instance for weeks, then they rebalance the load and the speeds invert. It's the hallmark of a scheduler managing pooled resources without user-visible affinity.
The "nicer waiting room" analogy is apt. The priority tier often just buys you a dedicated queue within the same overwhelmed processing facility, not a separate, faster facility.
—at
Welcome to the multitenant lottery. Your 45-minute add-to-list is the expected baseline, not an anomaly.
They optimize for batch throughput, not latency. That hamster wheel is a carefully balanced resource scheduler to keep their costs flat while selling you "priority" access to the same overbooked queue.
Ask for the average time-in-queue metric for your instance over the last week. They'll tell you it doesn't exist, which is your answer.
Prove it.
Exactly. The "priority" queue is a textbook example of non-work-conserving scheduling. It doesn't get its own resources; it just gets to cut the line in the shared resource pool. That's why your latency SLA is always a distribution, not a number. You're buying a better percentile in the same congested system.
If they actually gave you the average time-in-queue metric, you'd be able to model the entire M/M/c system. You could back-calculate their effective service rate and the number of virtual "servers" allocated to your tenant pool. The absence of that data is the control knob for their capacity planning.
numbers don't lie
Your point about testing with zero API steps is exactly what I do during vendor disputes. It removes their easiest deflection.
When they pushed the upgrade on us, I asked for the service-level objective for workflow queue *latency* versus *throughput*. They could only provide the latter, which confirms the hamster wheel analogy. They're built to move large batches slowly, not small actions quickly.
Your 45-minute benchmark during business hours is the key metric to hold them to. If you ever get an escalation, demand they reproduce that exact scenario in a support ticket and explain the queue position. They'll avoid it, but forcing the request creates a paper trail.
buyer beware, but buy smart