Used to get a reply on a support ticket within 4-6 business hours. That was acceptable.
Now? Average is 3-5 days. For a production model deployment issue last week, it took 7 days to get a first-line response. That's untenable.
The metrics:
* Ticket #44982 (API latency spike): Created Monday 10:00, first response Friday 15:00.
* Ticket #45110 (Feature store sync failure): Created Wednesday, first response the following Tuesday.
* Their public status page shows "all systems operational" while our tickets clearly indicate platform bugs.
This isn't a one-off. It's a trend. Funding should scale capability, not degrade it. Our team is now building more internal tooling as a contingency, which defeats the purpose of paying for their managed service.
Anyone else seeing this, or are we an outlier?
Prove it with a benchmark.
Yeah, we've seen similar delays. Our average first response is over 72 hours now. It's pushing us to document workarounds internally, which feels wrong for a premium service.
Do you think it's a support staffing issue, or are they just prioritizing larger accounts post-funding?
Yeah, those response times sound rough. We had a similar issue last month where a deployment pipeline was failing because of a change on their end, and it took four days just to get a "we're looking into it." By then we'd already rolled back and built a manual workaround.
It's frustrating when the status page says everything's fine but your tickets are clearly platform-related. Makes you wonder if their internal monitoring is missing something, or if they're just not linking support tickets back to wider incidents.
Did you guys notice if the quality of the initial response changed too? For us, the first reply feels more like a template now, asking for logs we already attached.
Learning by breaking
Those metrics are brutal, especially for production issues. The status page mismatch is what really gets me - if they aren't acknowledging platform-wide bugs, that makes the slow response even harder to stomach.
It feels like that funding round changed their priorities, and support got deprioritized. We're on a lower tier and I'm worried it'll only get worse.
Has anyone on a higher support plan seen better times, or is it bad across the board?
Yeah, the status page mismatch is a critical failure mode. It's not just a delay, it's a transparency breakdown. If their internal monitoring can't flag a platform bug generating multiple tickets, their SLA clock shouldn't start until they post an incident.
On your question about higher tiers: we're on their "Enterprise Critical" plan. Contractual SLA is 1 hour for P0. Last P0 (model serving 5xx errors) took 14 hours for a first *meaningful* response. They met the SLA with a 55-minute automated "we've received your ticket" email. So, it's bad across the board. The funding let them scale sales faster than support ops.
Benchmarks don't lie.
The status page mismatch you're seeing is a major red flag for operational maturity. When multiple customers are filing tickets for the same platform bug and it's not reflected as an incident, their internal communication between support and engineering is broken.
We hit a similar pattern at my last shop. The workaround tooling you're building now will become your permanent architecture if this continues. It's a costly distraction.
Have you considered escalating through your account manager? Sometimes finance pressure from a potential churn risk gets routed faster than a support ticket.
Based on the Enterprise Critical plan data point from user518, it's likely both. The funding probably created a staffing deficit as sales onboarded new large accounts, stretching the existing support team thin across a higher ticket volume. Simultaneously, the "first meaningful response" delay suggests they *are* prioritizing larger accounts, but the queue is so backed up that even those SLAs are effectively met with automation, not actual support capacity. The template replies asking for pre-attached logs are a classic symptom of an overwhelmed team using triage macros without reading the ticket.
CPU cycles matter
The SLA clock starting on automated receipt is a contract trick that only works once. After you get burned, you renegotiate the SLA terms to start on acknowledgment of an actual incident, not ticket submission. It costs more, but it's the only way to make the metric real.
Their ops scaling problem is classic. Sales gets funded, engineering builds features for new logos, and support becomes a cost center. The template replies are the canary in the coal mine for a team that's drowning in tickets they can't properly triage.
garbage in, garbage out
That contract trick is so frustrating. We got hit with it on our last renewal because I didn't know to look for it. Is "acknowledgment of an actual incident" usually defined by them posting to their status page, or something else in the contract wording? Trying to learn what to push for next time.
Still learning
That's a great question, and a tough one. In my experience, it's often defined by the first human response that confirms it's a legitimate issue on their side, not an automated receipt. Getting them to tie it to a status page update is ideal, but can be a hard sell. They might push back saying internal investigation is needed first.
You could try wording like "acknowledgment by a support engineer that the reported issue requires investigation." That moves it past the bot. Have you had any luck getting specific language like that into a contract before?
Those response times match what we've seen. The shift from 4-6 hours to multiple days for a first-line response is the critical failure.
The part about building internal tooling is the real cost. It's not just a delay, it's a permanent shift in your team's focus and architecture. Once that tooling exists, the ROI on their managed service plummets.
Escalate through your account manager now. Finance pressure is often the only thing that gets routed correctly when support ops are broken.
YAML all the things.
The permanent architectural drift is the hidden cost multiplier. I've seen teams build "temporary" circuit breakers and fallback logic that later dictates their entire deployment model, locking them out of the vendor's new features because they can't risk disabling their own mitigations.
On escalation, I agree that the account manager path is necessary, but it's only effective if you've quantified the business impact. Frame it in terms of their metrics. Tell them the delay is forcing your team to build "permanent contingency architecture" that reduces your platform's future consumption. That's a concrete revenue risk for their sales team, not just a support complaint.
No free lunch in cloud.