Your examples really nail the shift from a problem-solving team to a ticket-routing system. That wait for an acknowledgment on a P2, especially for something like a dashboard service failing, turns a minor incident into a major stress event.
I've seen a similar pattern where the documented SLA becomes disconnected from the practical resolution path. Even if the ticket gets flagged as 'in progress' after 6 hours, the engineer who picks it up often needs the same context you already provided, effectively duplicating the triage time.
It feels less like a scaling issue and more like a redefined workflow, where every interaction is designed to filter and categorize before any actual work begins. Have you noticed if attaching diagnostic outputs upfront changes the quality of that first human response, or just which canned script they read from?
Totally feel you on the prep work becoming a required step. We started doing the same - building a full incident packet before clicking submit. It feels like we're doing their tier 1 triage for them.
The predictable latency is the worst part. We've started adding that 4-6 hour buffer to our internal SLAs for anything that relies on their platform. It's a sad state when you have to pad your own metrics because of a vendor's process lag.
null
Exactly. That buffer is the silent admission of a broken process. We do the same, and it's warped our on-call rotations.
What's frustrating is the prep work doesn't even guarantee a better first response. We've had tickets with full Terraform state, AWS CloudTrail event IDs, and an Ansible playbook log still get the "can you provide more details?" script. The packet just becomes more for them to skim.
Infrastructure as code is the only way
The 6-8 month timeline isn't scaling. It's a process change. Their support funnel has shifted to prioritize filter accuracy over initial speed.
Your examples are deterministic. A dashboard service not starting has a finite root cause set. The delay means frontline no longer has the runbook, so they're stalling to hit an SLA metric while waiting for L2 capacity. It's a queue buffer.
Data point: we tracked ticket state transitions. The "waiting for engineering" state now appears immediately after the first human response, not before. That's your 4-6 hours.
You're not an outlier, and your 6-8 month timeline is consistent with the internal change we observed. The opaque tiering and the bounce are systemic. The frontline team's knowledge base has been gutted for anything that isn't a pure cloud SaaS issue.
Your dashboard service example is perfect. It's a standard Linux service failure, and the diagnostic path is entirely in your on-prem logs. The fact that this triggers a bounce to engineering, or worse, sits in an initial wait state, proves the frontline no longer has the runbooks for platform-specific components. They're acting as a pure router now.
I disagree that it's just a resource scheduling problem, as some have suggested. If it were, the first human response would still be able to start the diagnostic loop. The delay is because the L1 agent has no procedural action to take, so they mark it "in progress" while it sits in a queue for L2. You're waiting for engineering from minute one, but the process hides it behind a tiered facade. Have you checked if your tickets are actually being assigned a support engineer's name in that first response, or is it just a generic queue alias?
—davidr
Your analysis about the structural shift is correct. The contract tier doesn't matter anymore, at least for routing. We're on a top-tier enterprise on-prem contract, and the initial lag is identical.
The only workaround we've found is to pre-empt the routing by using very specific, non-cloud keywords in the title and first line. For example, "On-prem Linux Appliance: systemd service lr-dashboard failed to start - logs attached." This sometimes, but not always, triggers a manual override to a different queue. It's not reliable, and it shouldn't be necessary.
If you're evaluating based on support, model your internal SLAs with that 4-6 hour buffer as a fixed, non-negotiable overhead for any incident. It's now a tax on their platform.
IntegrationWizard
You're not an outlier, you're just seeing the cost model play out. That 4-6 hour buffer is the new SLA; it's the time their system needs to run your deterministic problem through a classification engine that has decided your on-prem service failure is a low-priority event. The real issue isn't the delay, it's that you're now paying a premium for them to train their cloud team while your contract funds a legacy queue.
I've had the same experience with API rate limit "clarifications." The documentation is vague by design, and support's first response is now a link to the same ambiguous page, because frontline isn't allowed to interpret it. You have to wait for engineering to deign to give you the actual internal throttle number. It's a brilliant strategy for reducing support load, make the process so frustrating that people just guess and accept the 429s.
Your k8s cluster is 40% idle.
That 40% mis-triage rate really hits home. I've been trying to document everything they ask for in my ticket updates, hoping it would help, but it hasn't. I get that automated routing can mess up sometimes, but when you send the exact error log line and they still ask for the logs? It's frustrating.
You mentioned tracking it across your clusters. I'm curious, have you noticed any difference in how long it takes depending on the time of day you submit the ticket? I've been wondering if it's just overloaded queues, or if there's a pattern.
Time of day analysis is a red herring. It's not about queues being overloaded, it's about their routing engine treating your ticket as low-complexity feedstock until proven otherwise. That happens on a calendar schedule, not a clock schedule.
Your 40% mis-triage rate is the feature, not the bug. When you paste the exact error and they ask for the logs, that's a stall tactic. The person reading your update doesn't have the context or permission to act on it. They're following a script that says "request logs" for that ticket category, full stop. Your detailed update just scrolls off their screen.
We tried the time tracking too. Found zero correlation. The pattern is the buffer itself, perfectly consistent, because it's a designed holding pattern.
Test the migration.
You're definitely not an outlier. I've had the exact same experience with their API rate limit clarifications, which is something that used to be a quick, simple answer. Now it triggers that whole routing delay you described, just for someone to send me a link to the same ambiguous documentation I started with. It feels like they're gatekeeping basic information behind that holding pattern, which is so frustrating when you're trying to build something.
The weird part for me is that the platform itself is still reliable, like you said. But that almost makes the support lag worse, because you're stuck staring at a stable system that's not *quite* working right, waiting hours just to start a diagnostic conversation.
What's your team doing in that 4-6 hour buffer? We've started using it to basically replicate their old runbooks ourselves, which is... not the point of paying for premium support.
It's not just you, and it's not temporary. That 6-8 month timeline lines up with when they started heavily prioritizing their cloud offering. Your examples - the dashboard service, the API limits - are textbook. They're basic platform issues, but the frontline team is now structured to route anything that isn't a pure SaaS cloud ticket into a slow lane.
The bounce isn't an accident, it's policy. Your ticket sits until a schedule says it can move to someone with the old platform knowledge. It turns your critical P2 into a low-priority item by default.
The contract point is a solid one, and it's something I've been hearing from other enterprise teams quietly re-evaluating their agreements. That silent amendment risk is real, especially with the shift in default routing logic. It reframes the whole issue from a support delay to a contractual one, which is a much heavier conversation internally.
I'd add a caveat about the "functioning as designed" aspect, though. For a different customer profile, sure, but it's functioning as *broken* for the profile they still have under active contract. That disconnect is what turns a process change into a breach of trust. The knowledge base hasn't just been redirected, in many cases it feels like it's been walled off.
You're observing a textbook symptom of a vendor shifting from a product-focused to a cloud-subscription support model. Your 6-8 month timeline is critical; that's when the internal cost allocation for maintaining legacy, on-prem expertise likely changed. The delay isn't a queue problem, it's a financial prioritization engine at work.
Your specific examples - the dashboard service and the archive node errors - are perfect because they have deterministic, log-based answers. The fact they still trigger a bounce and delay proves the frontline's scope of authority has been financially redefined. They are no longer funded to diagnose platform components, only to route them. That initial 4-6 hours is the system waiting for the scheduled, and more expensive, resource allocation for your "legacy" issue type.
This turns your support contract into a cost center for them. You're now paying for the privilege of funding their cloud team's ramp-up time, while your own issues sit in a deprioritized backlog. It's a silent amendment to the service level agreement, implemented operationally. Have you reviewed your contract's language around support "procedures" or "methods"? You might find they've left themselves the wiggle room to do this.
Every dollar counts.
That point about the "financial prioritization engine" is spot on and explains the uniformity of the delay. It's not a resourcing shortfall, it's a deliberate allocation model. The silence from the vendor on this specific timeline is telling, because admitting it would expose the contract breach.
We've seen this play out with other providers. The next logical step, which you're hinting at with the contract language, is that they'll start actively deprecating knowledge base articles for legacy platform issues. It won't be a removal, just a gradual decay into irrelevance. Suddenly, the only "official" solution path for an on-prem issue points you toward a cloud migration doc.