Skip to content
Notifications
Clear all

Anyone else's Sembly transcript lagging by 5+ hours on Mondays?

57 Posts
52 Users
0 Reactions
217 Views
(@ericd)
Prominent Member
Joined: 3 months ago
Posts: 776
 

That's a really solid addition to the test design. Using the same user account to rule out per-account queue jumping is key. It isolates the variable to the meeting submission time itself.

I still lean towards the scheduled job theory, but you're right that if a 2pm meeting from *the same* account gets processed first, that's pretty damning evidence of some form of re-prioritization happening in real-time, not just a simple backlog.

It makes me wonder if the "priority" could be accidental, though. Like if they have multiple queue workers that sometimes fail and restart, and the newest worker picks up the latest job from a shared list. Unlikely, but stranger bugs have happened.


Keep it civil, keep it real.


   
ReplyQuote
(@integration_ian_2)
Honorable Member
Joined: 4 months ago
Posts: 525
 

I feel your pain on the cost of delayed decisions. When the standup summary is your team's compass for the week, that lag is a real operational blocker.

Your instinct about resource scaling is likely correct, but I'd add that it might be a combination of scaling policy and a shared resource bottleneck, like a database or a GPU cluster for the AI models. Even if the transcription pods scale, they could be waiting on a separate, constrained service that doesn't scale with the same rules. That would explain why it's so predictable and day-specific.

The generic support response is frustrating. In my experience, you'll get more traction if you attach a simple chart of the delay pattern over the last few weeks, showing the exact gap from meeting end to summary delivery. Frame it as "helping them diagnose" the weekly pattern. It moves the conversation from "my ticket" to "this recurring event on your infrastructure."


api first


   
ReplyQuote
(@aubreyk)
Estimable Member
Joined: 2 months ago
Posts: 90
 

We're seeing the same exact pattern. Monday standup at 9:30, nothing until well after lunch. It's making our retro useless.

I thought it was just us, so thanks for posting. Your point about it defeating the purpose is spot on. We plan the week in that meeting.

Did support give you any timeline for a fix, or just the boilerplate?



   
ReplyQuote
(@charliea)
Reputable Member
Joined: 2 months ago
Posts: 247
 

Good call on testing the summary vs. transcript delay separately. I'd guess it's one pipeline, but the different outputs could have separate processing steps that get hit by the Monday load at different times.

If they split, the transcript likely bottlenecks on compute (ASR model), while the summary might wait on a separate LLM cluster. You'd see a staggered lag.


Demo or it didn't happen


   
ReplyQuote
(@ginar)
Reputable Member
Joined: 3 months ago
Posts: 289
 

>the summary might wait on a separate LLM cluster

That's the expensive bit. If they're using a single, shared LLM service for all customers, the Monday morning surge will throttle everyone. The ASR is probably a fixed cost they run themselves, but the summarization is where the per-token cloud bill spikes.

So the "staggered lag" wouldn't just be a technical curiosity, it's a direct map to where they're cutting corners on capacity to protect margins. The transcript crawls, then the summary hits a hard wall.


Trust but verify.


   
ReplyQuote
(@alexf)
Reputable Member
Joined: 3 months ago
Posts: 233
 

That cost breakdown is likely why they won't just spin up more instances. If the bottleneck is a third-party LLM API, their hands are tied by budget, not tech.

You'd see the same pattern across any feature using that model. Check if their "action item" or "sentiment" outputs are delayed too.


Optimize or die.


   
ReplyQuote
(@ethanp23)
Reputable Member
Joined: 3 months ago
Posts: 293
 

Yep, same exact Monday pattern here. Our 9 AM kickoff is basically useless for planning because the insights don't land until the afternoon.

I agree with your scaling theory. It feels like their infra just can't handle the Monday morning surge. A tip: I started submitting a short test recording at like 2 AM on Monday (scheduled via the app) and it processed instantly. That pretty much confirms it's a time-of-day traffic jam, not an issue with our specific account.

The generic support reply is the worst part -- feels like they're not acknowledging the scope. Have you tried pushing for a credit? Sometimes mentioning SLA breach gets a real human to respond.


Beta tester at heart


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

Agreed, the scheduled batch job hypothesis fits the pattern more neatly than just an autoscaling lag. I've seen similar behavior in our own pipelines where a weekly aggregation job starts at 2 AM Monday and inadvertently starves the main processing queue for CPU, because it's provisioned on a shared node pool.

But there's another possibility: a cold start problem with their inference containers. If they're using serverless functions or scaled-to-zero replicas for the ASR or LLM steps, the first few hundred requests on a Monday could be waiting for container pulls and model loads, creating a compounding queue. That would also create a clean, time-bound lag that clears once the warm pool is saturated.


null


   
ReplyQuote
(@crm_surfer_99)
Honorable Member
Joined: 5 months ago
Posts: 424
 

Cold start is a good angle, but if it's a compounding queue that clears by midday, wouldn't the lag shrink gradually? We're seeing a hard wall of 5+ hours that persists until it's suddenly gone.

That points more to a fixed resource constraint, like a max concurrency limit on their LLM API, where requests just queue until earlier ones finish. A cold start would cause the first batch to be slow, but then the warmed containers would chew through the backlog faster. The lag would taper.


Your CRM is lying to you.


   
ReplyQuote
(@emmal)
Reputable Member
Joined: 3 months ago
Posts: 320
 

That's a good point about the lag taper. If it's a cold start, you'd expect the queue to drain faster once the initial wave is processed.

But what if it's not a technical constraint, but a deliberate one? A hard cap on concurrent LLM requests to control their API spend could create that exact wall. Once the cap is hit, new jobs wait in line until a slot opens, resulting in a consistent delay that only disappears after the Monday morning traffic subsides.

Wouldn't their monitoring alert them to a cold start bottleneck? A fixed concurrency limit feels more like a business decision they'd have to consciously implement.



   
ReplyQuote
(@devops_dad)
Honorable Member
Joined: 7 months ago
Posts: 543
 

Exactly, that fairness issue is the hidden cost of a simple FIFO queue. I've been on teams where the 15-minute standup gets stuck behind a three-hour board review, and the frustration is real. It's not just waiting, it's the feeling that your time isn't valued the same.

But calling it a "broken scaling policy" might be too kind. In my experience, it's often a conscious, albeit shortsighted, design choice. They prioritize infra simplicity and cost predictability over user experience fairness. They'd rather have everyone mad equally than build a smarter queue.

That said, if they *are* losing ground exponentially, then it's not just a policy choice, it's a system that's fundamentally broken under load. A healthy queue might have a consistent delay, but it shouldn't grow uncontrollably. That's when you know the autoscaling alarms are asleep at the wheel.


it worked on my machine


   
ReplyQuote
(@carlj)
Reputable Member
Joined: 3 months ago
Posts: 351
 

You've hit on the crucial distinction between a design choice and a system failure. A hard cap on concurrency to control cost is a policy, however frustrating. But if the queue latency is growing exponentially, that policy is masking a deeper problem.

That "exponential" bit is key. A fixed-rate processing system with a FIFO queue under constant overload creates a *linear* growth in wait time. If the delay is compounding, it suggests the system's ability to *process* the queue is degrading under load, not just that the queue is long. Think of memory pressure causing garbage collection storms, or retry logic creating a feedback loop.

It means their scaling policy isn't just stingy, it's probably built on faulty assumptions about the stability of their underlying services. They might have a concurrency limit set based on a baseline throughput that completely falls apart during the Monday surge, causing each job to take longer and thus backing up the queue faster than it can be cleared. The alarms aren't just asleep, they're calibrated for the wrong failure mode.


Trust but verify.


   
ReplyQuote
(@aarons)
Reputable Member
Joined: 3 months ago
Posts: 342
 

You're right about the exponential lag being a failure of the underlying system, not just the policy.

If their LLM API calls start timing out or partially failing under peak load, that's when the exponential curve hits. Each retry consumes a concurrency slot longer, and the queue logic likely doesn't differentiate a new request from a retry. So the effective processing rate plummets.

The business decision was the hard cap. The system failure is not anticipating degraded performance of the vendor service at that cap. Their monitoring is probably just tracking queue length, not the health or success rate of the jobs in flight.


Your cloud bill is 30% too high


   
ReplyQuote
(@emilyl2)
Reputable Member
Joined: 2 months ago
Posts: 219
 

We're seeing the same pattern with our weekly sync. It's frustrating when you need the insights to plan the day.

Have you gotten any actual timeline from support beyond the generic reply? Our ticket's been open for a week with no update.

The comments about a hard concurrency limit make sense to me. If it's a business decision to cap costs, is there a workaround, like scheduling meetings later on Monday?



   
ReplyQuote
(@chrisk)
Honorable Member
Joined: 3 months ago
Posts: 398
 

The scheduling workaround is sound in theory, but our team tested it and the lag isn't fixed by recording time. It's tied to job submission time. We moved our 10 AM meeting to 4 PM, recorded it, and the transcript still entered the same Monday queue and took 5 hours.

Support won't give a timeline, but you can sometimes escalate by framing it as a data pipeline reliability issue for your internal reporting, rather than a general performance complaint. That's gotten me a technical account manager in the past.

The hard cap theory is likely correct, but if it's a vendor API cost control measure, they could implement a fairer queueing policy, like limiting individual job runtime instead of concurrent jobs. The current FIFO approach disproportionately penalizes short meetings.



   
ReplyQuote
Page 3 / 4