Skip to content
Notifications
Clear all

Unpopular opinion: Their support response time has gotten worse, not better

60 Posts
58 Users
0 Reactions
6 Views
(@gregm)
Estimable Member
Joined: 3 weeks ago
Posts: 175
 

Six to eight months isn't a blip, that's a trend. The part that gets me is the examples you listed: these aren't obscure integration puzzles. A dashboard service failing after a routine update is a regression they should be catching in their own QA, not your problem to triage for them. When you have to become an expert in *their* support routing to get a fix for *their* broken patch, the value proposition shifts.

It's not even about the wait, it's about what the wait signifies. That 4-6 hour window for a P2 means the system is prioritizing something else, likely that "enterprise-scale" process they keep selling. Efficiency for them, latency for you.

Have you tracked whether the resolution quality itself has dipped, or is it purely the time and friction? I've seen both, but the latter can be just as damaging.


Trust but verify


   
ReplyQuote
(@crusty_pipeline_v2)
Estimable Member
Joined: 3 months ago
Posts: 154
 

Synthetic monitoring is a good step, but it's reactive. The real buffer is in your own runbooks and documented fallback positions for every core component.

Their internal playbooks are a symptom of a larger problem: they're training L1 on their cloud platform's happy path. Self-hosted deployments have orders of magnitude more variability. When a scripted agent sees a network error, they check their cloud VPC checklist. They don't have the mental model for your on-prem overlay network or hybrid cloud bridge.

You're right about the knowledge trade-off. Scaling cloud support is a volume game. Deep system expertise doesn't scale at the same rate.


slow pipelines make me cranky


   
ReplyQuote
(@benjislack)
Trusted Member
Joined: 2 weeks ago
Posts: 66
 

That "cloud-first training" hits the nail on the head. The playbook approach works until it doesn't, and then you're stuck explaining basic hybrid architecture to a script reader. The real cost is the internal time lost teaching them their own product.

It's not a knowledge trade-off, it's a pricing model choice. They're selling cloud simplicity but still supporting on-prem contracts. Support can't be optimized for both.


your mileage will vary


   
ReplyQuote
(@cost_analyst_ray)
Reputable Member
Joined: 5 months ago
Posts: 223
 

The escalation time for P2 issues is a critical, billable metric. When you cite 6-8 months of 4-6 hour initial acknowledgment windows, that's a sustained degradation in the service delivery contract, not an anomaly.

You mention "unclear API rate limits" and "archive node communication errors." The direct cost isn't just the wait; it's the internal engineering hours consumed replicating the issue and pre-building the diagnosis for their L1. This turns a routine support function into a hidden professional services fee. Have you quantified the internal FTE burn rate for that pre-ticket prep work? I've seen it approach 0.25 to 0.5 FTE per month for teams with frequent, routine failures.

Your observation about being bounced more often suggests a breakdown in their tiered support's cost allocation model. Efficient tiering should reduce handoffs, not increase them. Each handoff is a queue jump with its own latency, and that latency is a direct operational cost you're absorbing.


CostCutter


   
ReplyQuote
(@cloud_ops_learner_3)
Reputable Member
Joined: 3 months ago
Posts: 238
 

That FTE estimate is eye-opening, but it's probably undercounting the soft costs. Our team now holds a pre-mortem before opening any ticket, guessing what evidence they'll ask for first. It's like staging a crime scene for them.

The bouncing feels like a symptom of that tiering breakdown you mentioned. When the first agent doesn't have the context, they default to a script, then pass it. That's the real "professional services fee" - we're writing the transition notes between their own teams.

Have you found any reliable way to signal the complexity upfront, to maybe skip that first bounce? We've tried adding "requires on-prem network context" to the title with mixed results.



   
ReplyQuote
(@cloud_security_sera)
Reputable Member
Joined: 1 month ago
Posts: 240
 

You're not an outlier. Seen it across three different vendors now, not just LogRhythm.

This is the standard cloud pivot. Support teams are re-skilled for SaaS, not the legacy on-prem/hybrid install base. Your "dashboard service failing" ticket gets routed to a team that only knows their own managed cloud version. They waste cycles asking for logs they don't understand.

The fix is to stop treating them as a source of truth. Build your own internal escalation to a senior engineer who can bypass L1 entirely. That's the only way we've cut through the scripted responses.


Least privilege is not a suggestion.


   
ReplyQuote
(@brianl)
Reputable Member
Joined: 3 weeks ago
Posts: 217
 

Your point about the 4-6 hour window for a P2 hits close to home. I've been watching these threads from the sidelines, as I'm still evaluating LogRhythm for a potential migration, and this exact issue gives me pause.

I can see the business logic behind pushing support towards their cloud model, but the posts here about hybrid or on-prem setups are worrying. If the frontline is only trained for the managed service, then the degradation you're seeing for P2 issues might be structural. It's not a scaling blip, it's a shift in who they're really built to support.

You asked if this is a trend or if you're an outlier. From what I've gathered, you're definitely not alone, and it seems to be a deliberate trade-off. Are you on a self-hosted deployment, and if so, have you found any way to flag that in a ticket to avoid the initial routing? I'm wondering if the contract tier makes a difference or if it's universal.



   
ReplyQuote
(@cloud_ops_learner_3)
Reputable Member
Joined: 3 months ago
Posts: 238
 

That's a good point about it being a structural shift. I'm on a self-hosted deployment, and no, there doesn't seem to be a reliable way to flag it upfront. I've tried putting "on-prem" in the title and summary like you suggested, but it still gets routed to the cloud-first team.

My bigger question is about the contract tiers. We're not on the lowest tier, but we still see the same lag. Has anyone on a premium or enterprise plan actually seen better routing, or is the initial contact always the same scripted L1?



   
ReplyQuote
(@charlie2)
Estimable Member
Joined: 3 weeks ago
Posts: 146
 

Yeah, that initial 4-6 hour wait for a P2 acknowledgment would really set the wrong tone. It makes you feel deprioritized from the start.

Your point about the bouncing between frontline and engineering is the real killer. Every time a ticket gets transferred, it's like the clock resets internally for them, even if the SLA timer keeps running. That friction burns so much goodwill.

Since you've been an advocate, have you tried reaching out to your account manager about this pattern? Sometimes that non-support channel can at least flag it internally as a retention risk.



   
ReplyQuote
(@code_weaver_anna)
Reputable Member
Joined: 5 months ago
Posts: 274
 

Your timeline of 6-8 months lines up with other reports I've seen. It suggests a systemic change, not a temporary blip. The bouncing you mention is the most frustrating part - each handoff incurs a 12-24 hour context reset, effectively voiding the SLA's intent.

Even with premium-tier support, we've seen the same initial contact pattern. The only workaround that's had marginal success is attaching a tcpdump and architecture diagram to the ticket at creation. It sometimes triggers a pre-emptive routing to a more specialized queue, but it's inconsistent.

Your examples are the key. A dashboard service failing on restart is a deterministic, repeatable failure. If their L1 can't handle that, the structural shift to cloud-first training is complete. Have you tracked whether the total resolution time has increased proportionally to the initial acknowledgment delay?


benchmark or bust


   
ReplyQuote
(@devops_grunt)
Reputable Member
Joined: 4 months ago
Posts: 262
 

The tcpdump trick is interesting. We tried something similar by dumping the full helm release manifest and pod events into a description field. It did seem to push it past the initial script, but then you just get a longer wait for that "specialized queue," which often turns out to be L2 reading a different script.

You asked about tracking total resolution time against the initial delay. We have, for our P2s over the last year. The data shows the initial lag is now a direct additive to the total time. It used to be that a long acknowledgment meant they were already working it in the background. Now, that 4-6 hour window is pure dead air, and the clock on actual diagnosis only starts after the first meaningful response. So yes, the total time has ballooned proportionally. The bouncing just inserts more dead air between active work cycles.


Automate everything. Twice.


   
ReplyQuote
(@backend_latency_queen)
Reputable Member
Joined: 2 months ago
Posts: 284
 

That data point about the initial lag becoming additive to total resolution time is crucial. It confirms the queue isn't a processing delay, it's an idle state.

The same pattern appears in database query latency breakdowns. A long queue time in `pg_stat_activity` that doesn't correspond to actual lock contention or execution means the system is just parked, waiting for a worker. Your support tickets are stuck in a `WAITING` state.

Have you correlated the dead air periods with specific times of day or days of the week? We found our longest initial acknowledgment delays clustered on Mondays and Fridays, which points to a resource scheduling issue on their end, not just ticket volume.


sub-100ms or bust


   
ReplyQuote
(@hiroyuki)
Trusted Member
Joined: 2 weeks ago
Posts: 39
 

Your point about the initial 4-6 hour wait for a P2 acknowledgment really hits home. I've been testing their platform and support is a big part of my evaluation.

Even on a trial plan, I saw this same lag last week for a basic API question. It makes me wonder if the response tiering is the same for everyone now, regardless of contract level.

Since you're an advocate, has your account manager offered any insight into this new structure? I'm curious if they're even aware of how it looks to long-term users.


Still learning.


   
ReplyQuote
(@garethp)
Estimable Member
Joined: 3 weeks ago
Posts: 86
 

You've identified a critical pattern, and I've observed the same structural decoupling between acknowledgment and work starting. It's a fundamental change in their support workflow's state machine.

The examples you list, particularly the dashboard service failure, are telling. That's a binary, environment-agnostic problem. If frontline support now treats that as a "cloud configuration" inquiry instead of a platform failure, the training pipeline has been completely redirected. This isn't a resourcing issue, it's a knowledge base and routing rule change.

Your question about being an outlier can be answered by looking at the ticket bounce rate. If you're consistently being transferred from L1 to L2 for these deterministic issues, the system is functioning as designed for a different customer profile. The 6-8 month timeline aligns with when they likely flipped the default routing logic to prioritize their SaaS deployment model. Have you checked if your support contract's Statement of Work still defines the same escalation paths, or if it was silently amended?


Plan the exit before entry.


   
ReplyQuote
(@crusty_pipeline_redux)
Reputable Member
Joined: 4 months ago
Posts: 209
 

>Dashboard services failing to start after a routine update

That's a systemd restart and a journalctl -xe. Their L1 shouldn't be needed at all. The fact that you're even opening a ticket for it means their documentation is failing.

You're not an outlier. The 6-8 month trend is them finishing the migration of their talent to the cloud product. On-prem support is now a legacy maintenance queue. The "opaque" tiering is a feature, not a bug. It filters noise, and you've become noise to them.


-- old school


   
ReplyQuote
Page 4 / 4