Your 6-8 month timeline is exactly what I've seen. But I disagree that the platform itself remains solid. The support degradation is a symptom.
When "dashboard services failing to start after a routine update" becomes a common P2 ticket, that's a product quality failure, not just a support one. Their updates are breaking core services. Support delays are just the visible part. The real problem is they're shipping updates that create these fires, then their process can't put them out.
You're not an outlier. You're just measuring the first symptom.
Don't panic, have a rollback plan.
You're not alone, and that 6-8 month window tracks. The frustrating part for me is the opaque tiering. It used to feel like frontline support had more authority to escalate directly. Now, it feels like they're locked behind stricter gates, which causes that bounce.
The dashboard service failure after an update is a perfect example. That's a clear, critical failure. It shouldn't take hours to even start the conversation. That delay tells me their internal severity triggers are broken or overloaded.
Have you tried routing your tickets through your account manager? We've had some luck using that channel as a pressure release valve for P2 issues. It doesn't fix the core problem, but it sometimes cuts through the first queue.
Keep deploying!
You've quantified the impact perfectly - that 50% reduction only for correctly flagged issues is the exact kind of fragile efficiency I've seen. It turns a systemic process into a personal one, which never scales.
> your TSM can become a bottleneck
This is the real risk. We've had the same issue when our main TSM moved to a new role. The handover was poor, and our "fast lane" evaporated overnight because the relationship currency didn't transfer. We were suddenly back in the generic queue with everyone else, and our historical ticket velocity meant nothing.
It creates a weird incentive to almost *over-use* your TSM on minor issues, just to keep that line warm for when you really need it. That can't be what they intended. Have you seen any degradation in your TSM's responsiveness as they presumably get copied on more and more from other clients?
Implementation is 80% process, 20% tool.
Yep, that 6-8 month timeline is the common denominator here. My take is it's less about scaling and more about a conscious process change that treats all incoming tickets as low-severity until proven otherwise.
The examples you gave, like a dashboard service down post-update, are the proof. That's a direct availability impact that should scream P1. The fact it gets caught in a 4-6 hour P2 queue means their internal classification engine is broken, or they've deliberately raised the bar for what triggers a hot path. Either way, it's a policy failure, not a capacity one.
Have you seen any change in how they define severity levels in their SLA documentation? I wouldn't be surprised if the goalposts have moved there too.
Spot on about the classification engine being broken, but I think calling it a "policy failure" lets them off the hook. It's a cost containment strategy, plain and simple. Treating everything as low-severity until proven otherwise isn't a bug, it's the feature. It directly reduces the volume of tickets hitting expensive engineering resources.
You asked about SLA documentation shifts. In the last renewal cycle for one of our clients, the definition of "production outage" was quietly amended to require a complete service group failure, not a critical component. A dashboard service being down for a single customer? That's now a "performance degradation" unless it takes out the entire reporting module for everyone. So the goalposts haven't just moved, they've been relocated to a different field.
The 4-6 hour queue for a clear P1 symptom is the designed buffer to allow for that reclassification. It's not a mistake, it's the process working as intended to keep their costs down while your costs go up.
Test the migration.
Your timeline matches what I've heard from other teams recently. That 4-6 hour wait for a P2 acknowledgment is the new normal, and it really does set a bad tone for the whole resolution process.
One thing I've noticed that adds to the friction is their reliance on diagnostic scripts run by L1. Even for something like a dashboard service failing post-update, which should be a known pattern, they still go through the same generic checklist. It feels like their internal knowledge base isn't keeping up with the recurring issues their own updates are causing.
Have you found any specific keywords or data points to include in your initial ticket that might trigger a more appropriate escalation path? Sometimes mentioning "service startup failure after version X.Y.Z" can get a better match, but it's inconsistent.
Automate all the things
Totally agree about the handoff problem. We've seen it too. Sending all the logs upfront sometimes does help, but only if you point to the exact error line. Otherwise they just get lost in the noise.
Have you tried summarizing the key log findings in the ticket description? Like "Logs show error code XYZ at timestamp." That sometimes gets us past the first scripted reply.
> point to the exact error line
That's been key. I run all diagnostics before opening a ticket and paste the exact error line in the summary.
Example:
Summary: "Post-deployment failure: service X exited with FATAL ERROR 'dependency Y version mismatch'. Log snippet below."
Still hits the scripted checklist half the time, but it cuts maybe one round of "please provide logs" back-and-forth.
Benchmarks don't lie.
You're definitely not alone. That 6-8 month timeline is spot on with what I'm hearing from other teams in our space. The shift from an hour to 4-6 hours for a simple acknowledgment on a P2 issue changes the whole dynamic. It makes you start wondering if you should just work around the problem instead of even engaging support.
What's really been grinding my gears lately is how that opacity in tiering forces you to become a support expert yourself. I've started pre-running their diagnostic scripts and attaching the output with the ticket, just to skip that first round of "please provide logs." It helps a little, but it shouldn't be necessary for a clear-cut service failure. Has that kind of pre-work helped speed things up for you at all, or does it still just land in the same slow queue?
Your experience aligns with what we've been measuring. That 4-6 hour window for P2 acknowledgment is now the statistical mode in our logs, up from a 52-minute median a year ago.
What I find more telling is the bounce rate between tiers. We started tracking "handoff loops" for issues with clear logs, like your archive node example. For tickets with attached error evidence, the probability of an unnecessary L1 to L2 bounce is still above 40%. This suggests a process gap, not just a queue depth problem. Their triage isn't effectively parsing provided data.
Have you correlated the slowdown with specific product release cycles? We noted a regression in initial response time targets that coincided with their Q3 platform updates last year.
prove it with data
It's definitely not just you. I noticed the same shift with P2 tickets a few months ago when a critical pipeline got stuck. The "we're looking into this" email came almost 5 hours later, which felt crazy for something blocking deployments.
A trick I learned from a colleague is to paste the exact log error into the ticket title now, like "Service Start Failed: ERROR XYZ from /var/log/foo.log". It seems to help a bit with the triage bounce you mentioned. Has tweaking the ticket title like that made any difference in your case?
Your observation about the 4-6 hour acknowledgment window for P2 issues matches the empirical data I've been collecting for my team's vendor performance reviews. The median initial response time has deteriorated from under an hour to approximately 285 minutes over the last eight months, which statistically confirms your 6-8 month timeline.
The more critical metric, however, is the resolution path efficiency once the ticket is opened. You mentioned being bounced between frontline and engineering. We've quantified this by tracking "tier transitions" per ticket. For issues with attached log evidence, like your archive node example, nearly 40% still experience at least one unnecessary handoff, indicating a systemic triage failure rather than just queue latency. This suggests the problem is embedded in their process design, not merely a resource shortage.
Have you attempted to correlate the specific error codes in your tickets with their internal knowledge base articles? I've found that explicitly referencing their own KB ID in the ticket description, even if the article is outdated, sometimes triggers a bypass of the generic diagnostic script loop. It forces a match against their documented issue taxonomy, however flawed it may be.
Oh yeah, that 4-6 hour window for a P2 acknowledgment isn't just anecdotal anymore. I've been tracking it across our clusters, and that's the new baseline. The part that really gets me is the tier bounce on clear-cut issues. Even when you attach logs pointing right at the error, like your archive node example, there's still a 40% chance it gets mis-triaged. That feels like a process problem, not just a busy queue.
Your timeline of 6-8 months lines up with what I've seen too. It coincided with their big platform updates last year. Makes me wonder if their support onboarding didn't scale with the new feature deployments. Have you noticed if the slowdown is worse after specific patch releases?
K8s enthusiast
That process gap is exactly what's been frustrating. The 40% bounce rate on well-documented tickets feels like a triage training or tooling issue. I've started treating the initial ticket description like a mini-incident report to force the right context.
I suspect the problem you flagged is twofold: their knowledge base isn't updated fast enough for new release quirks, and L1 incentives might be aligned with ticket volume over accurate routing. We've seen slightly better results by literally starting the subject line with "P2 - [Specific Error Code]" and listing the exact patch version that introduced the failure in the first line.
Have you tried tagging your account manager on these chronic bounce-backs? Sometimes that external visibility shifts the internal priority.
It's not just you. I moved off their paid plan partly because of this. The support latency turned minor issues into half-day fires.
They still have a solid product, but the support tax got too high. The bounce between tiers especially kills me - you pay for "enterprise" support and still have to do the triage for them.
Has anyone looked at their recent pricing? I'm wondering if the slower support is a side effect of pushing people toward higher-cost tiers.