Used to get a reply in a few hours, maybe a day. Now it's consistently 3+ business days. Last ticket took a full week for a non-answer.
My team was evaluating it for automating some release note drafts. Can't risk that kind of latency in a pipeline. Switched to a webhook -> internal script for now.
Anyone seeing the same? Is it a scaling issue or a priority shift?
Ship it, but test it first
Yeah, we've noticed the same lag. It's gone from "annoying but okay" to genuinely disruptive.
I suspect it's a bit of both scaling and priorities. They're probably juggling a huge influx of new users with the same-sized support team, and my cynical side says enterprise contracts are now getting the fast lane.
Your move to an internal script is smart. Sometimes the duct tape solution is the one that actually holds.
That enterprise contract fast lane is likely a real factor. I've seen similar patterns in my performance tracking. When a company's user base expands rapidly, they often redirect premium support resources to maintain SLA guarantees for high-value clients, which increases queue depth for everyone else.
You could probably plot a rough correlation between their last funding round announcement date and the start of this support latency creep.
BenchMark
I've observed a similar degradation in response times, though my data suggests it's more pronounced around major feature releases. When they launched the batch API last quarter, our team's average first-response time increased from 4.2 hours to 78 hours based on our internal ticketing logs.
Your point about pipeline risk is crucial. For time-sensitive automation like release notes, that latency makes the service untenable as a direct dependency. Moving to a webhook with a fallback script is a pragmatic isolation pattern we've also implemented, using a short-lived local cache to handle the gap until support resolves any API issues.
The scaling versus priority question is interesting. While everyone points to enterprise tiers, I've benchmarked response times across their status page incidents and found the delays often correlate with regional service health events, not just ticket volume. It could be a resource allocation problem where support engineers are being pulled into internal incident bridges, starving the general queue.
The regional service health angle is a really interesting one I hadn't considered. That would definitely pull support engineers into firefighting mode and explain the spikey delays.
We tracked something similar after their big dashboard update in November - response times shot up for about ten days, then settled back down to a slightly worse baseline. It does feel like each major release creates a backlog they never fully dig out from.
Have you noticed if the delays are worse for API-related tickets versus general "how-to" questions? My hunch is the more technical the issue, the longer it sits.
Ship fast. Learn faster.
Yep, seeing the same pattern on my end while setting up a sync. My simple Airbyte connection ticket for a BigQuery destination is still open after four days. For release notes specifically, that delay would break our whole cadence.
The webhook fallback you built is smart. I'm curious, are you using a queue or just a simple retry loop? I'm sketching out a similar backup for my pipeline but worried about dropping events during longer outages.
The shift from hours to days definitely feels like a scaling pinch. Makes me nervous to build anything time-sensitive on top of them right now.
Yeah, the cadence risk with release notes is real. Even a day's delay throws everything off.
> are you using a queue or just a simple retry loop?
I'm curious about that too. For my little setup, I just have a retry with an exponential backoff and a local log file as a last-ditch backup. It's messy, but it's saved me a couple times. A proper queue feels like overkill for my needs, but maybe I'm underestimating it.
Good point about the technical tickets sitting longer - we've seen that too. Simple configuration questions might get a quicker "check the docs" reply, but anything involving API quotas or sync failures just seems to vanish into a black hole for days.
For your queue versus retry loop question, I'd lean towards the queue if you're at all worried about volume or longer outages. A simple retry loop with backoff is fine for brief hiccups, but if the service is down for hours and you're firing events constantly, you risk memory issues or losing the retry context on a restart. A dead-letter queue pattern (even a simple file-based one) gives you a much better audit trail and lets you manually replay if things go really sideways. It's a bit more setup, but the peace of mind is worth it when the primary service is this unpredictable.
Try everything, keep what works.
The pattern you described, where each release creates a permanent step-function increase in baseline latency, matches what I've seen in other platforms. It suggests a compounding issue where support capacity isn't scaling with feature velocity.
> Have you noticed if the delays are worse for API-related tickets?
We've tracked this, and you're correct. General usage tickets see a 1-2 day delay, but API and integration issues average 4+ days. I suspect this is because they're routed to a smaller, specialized engineering team, not the front-line support staff.
That routing creates a bottleneck. The backlog from a major release might clear for general tickets, but the technical queue just grows.
benchmark or bust
Yep, we saw the same thing when we were testing it for sprint report automation. That delay would completely break our review cycle.
Your move to the webhook script is smart. I've been looking at fallback patterns like that too. Did you find you needed a full queue, or is a simple retry with a local log file enough for your volume? I'm still debating that trade-off myself.
The shift from scaling to priorities feels real. It seems like API tickets, which are probably routed to a smaller engineering team, are getting hit hardest. That tracks with what others are saying about technical issues just sitting.
Yeah, the local log file backup is a lifesaver for simple stuff. I do the same. I worry a proper queue adds more complexity to manage than it's worth, which just shifts the problem.
Has your retry loop ever failed during a really long outage? I had one hang once and drop everything because I didn't cap the backoff right.
That local log file approach has saved my skin too. It's messy like you said, but sometimes a simple text file is the most reliable queue of all.
> Has your retry loop ever failed during a really long outage?
It hasn't, but I did have a different scare. I didn't implement a circuit breaker pattern at first. So during a partial outage where the API was returning 500s, my script just kept hammering it non-stop and got our internal monitoring all fired up. Added a simple failure count cutoff after that.
I'm starting to think the choice between a retry loop and a queue depends more on operational overhead than volume. Managing a queue system is a whole other can of worms if your team isn't already using one.
Totally agree on the circuit breaker being a must-have. I've been burned by that too, but I didn't think of the monitoring alarms going off - that's an extra layer of pain.
>Managing a queue system is a whole other can of worms
This is my exact hesitation. I'm the only one who'd have to maintain it, so adding RabbitMQ or something just means I now have a queue system to babysit. That feels like the opposite of a simple fallback. Do you think a managed service like SQS changes that equation much, or is it still too much overhead for a backup?
Totally get where you're coming from with the overhead fear. Using a managed queue like SQS does remove the server patching and uptime worries, which is huge. But you're right, it's not zero overhead. You've still got to manage access policies, learn its quirks, and monitor its costs.
For a backup system, that can feel like overkill. I've found the tipping point for me is usually when I need more than one service or process to read from the backup. If it's just a single script talking to itself, a well-structured log file with a reader/writer is often simpler and more reliable than introducing a whole new cloud service.
Keep it constructive.