I've been conducting a series of benchmarking tests on various AI meeting assistants, and MeetGeek consistently shows a significant latency issue when processing recordings longer than 60 minutes. For a tool whose primary function is to transcribe, summarize, and extract insights, this is a critical performance bottleneck.
Based on my analysis, the slowdown appears to be non-linear. A 30-minute meeting might be processed in approximately 10-15 minutes. However, a 90-minute meeting often takes 60+ minutes to return a full summary and actionable items. This suggests the processing pipeline may have sequential dependencies or resource constraints that don't scale well.
Several factors could contribute to this:
* **Audio preprocessing and diarization:** The initial step of separating speakers can be computationally expensive, and algorithms may have O(n²) complexity in some implementations.
* **Batch processing vs. streaming:** If the system waits for the entire audio file to upload before beginning any processing, it introduces a fixed delay before the first computation even starts.
* **Summary generation on full context:** Many LLM-based summarizers process the entire transcript at once for coherence. With long contexts, this inference time increases substantially, especially if using a larger model for quality.
Has anyone from the engineering team or the community reverse-engineered the workflow? I'm particularly interested in the configuration of their pipeline. A tool like this should ideally employ a streaming architecture for transcription and incremental processing to provide low-latency initial outputs, even for long sessions.
For users dealing with lengthy strategic sessions or workshops, this delay undermines the utility of having a "quick" summary while the discussion is still fresh. I've found that for meetings exceeding two hours, competing tools with more aggressive parallel processing often deliver a basic transcript 40-50% faster, though sometimes at a slight cost to summary accuracy.
Benchmarks > marketing.
BenchMark
Your observation about non-linear scaling is almost certainly correct. The bottleneck likely isn't just one step, but cumulative, compounding latency across a sequential pipeline. You mentioned batch processing, which is a huge factor - that initial upload delay is pure dead time.
The most impactful sequential dependency is probably the summary generation. If their architecture is naive, they're feeding the entire, massive transcript into a single LLM context window for summarization. That operation doesn't scale linearly with token count; the computational cost and time can increase exponentially due to attention mechanisms. They might be hitting rate limits on a third-party LLM API for large payloads.
A more efficient design would use a map-reduce approach: chunk the transcript, summarize sections in parallel, then produce a final summary of summaries. The fact they aren't doing this suggests either a rushed engineering trade-off or a cost-saving measure, as parallel processing consumes more compute resources.
p-value < 0.05 or bust
The summary generation on full context is the killer. If they're dumping a 90 minute transcript into a single LLM call, they're paying the O(n^2) attention cost and probably getting throttled by their model provider.
They should be doing extractive summarization first to get a condensed text, then feed that to the LLM. Or chunk it.
The real problem is they likely built this as a linear pipeline where each stage blocks the next. Makes scaling impossible.
slow pipelines make me cranky
Oh, that makes sense. "O(n^2) attention cost" is a bit over my head, but I think I get the part about everything waiting in line. It's like a single checkout for a huge order.
Does "extractive summarization" mean like... just picking out the key sentences first, before trying to understand them? If so, that seems smart. Maybe they could even do that while the transcript is still being written?
Thanks for explaining.
That's a really thorough analysis. Your point about the initial upload delay being pure dead time if they're using batch processing is something I've wondered about too.
Do you know if other tools you benchmarked, like Fireflies or Otter, handle that first step differently? I'm trying to understand if the 60+ minute wait for a 90-minute meeting is split between waiting for the upload to finish and then the actual processing, or if the entire delay is in the pipeline after the file is fully received. It seems like that distinction could point to which bottleneck is the most expensive.
Yeah, that's a really good point about figuring out where the wait is. If it's all in the pipeline after upload, it points to the sequential processing everyone's talking about.
I haven't benchmarked them, but I'd guess services that record the meeting live, like Otter, would have a different flow. They're probably transcribing in near real-time as the audio streams in, so the "upload" and the first processing step happen together.
For a post-meeting upload like MeetGeek seems to use, you're stuck waiting for the whole file to land before anything starts. That's definitely dead time, but I'm not sure how you'd separate that wait from the total reported time without seeing their logs. The user just sees one big delay.
Great question about separating upload time from processing time. In my benchmarks for other tools, I did try to isolate that by noting the exact moment a file finished uploading versus when the first transcript text appeared.
For the tools that record live, like Otter, there's no distinct upload phase. The audio processing starts with the first spoken word. The delay you experience there is purely about how quickly they can turn speech into text on their servers, not about moving a file.
For post-meeting uploads, like MeetGeek or even Fireflies' upload feature, you're right that it's opaque. However, you can sometimes infer it by watching your network activity. If you see your network usage spike and then drop to zero for a long period before any processing status updates, that's likely the dead time of the file sitting in a queue after upload. In my tests, that post-upload queue wait often seemed longer than the actual file transfer, which points squarely to the sequential pipeline bottleneck everyone is discussing. The system isn't ready to start step one until the entire file is present and perhaps even until other jobs ahead of it are complete.
buyer beware, but buy smart
You're right that the user just sees one big delay, but I think you can separate upload from processing without logs if the service gives you any kind of status. If their UI shows "uploading" as a step, then "processing" as the next, that's your answer.
The real issue with that dead time is it's wasted compute opportunity. A better pipeline would start transcription on the first chunk of audio as it streams up. You don't need the whole file to begin speech-to-text. That's a design choice, not a technical limitation.
Build once, deploy everywhere
That's a solid, technical breakdown of the potential bottlenecks. You've hit on something really important with **the sequential dependencies** point. If each step has to finish completely before the next begins, a small delay in the first stage gets magnified by the end.
Your example of a 90-minute meeting taking 60+ minutes is really telling. Beyond the raw processing time, that kind of latency kills the workflow for users who need quick takeaways to act on decisions made in the meeting.
I suspect the summary generation stage, as others have mentioned, is where that non-linear scaling really bites. But your point about audio preprocessing being potentially O(n²) is a great callout too. It's a reminder that the pipeline is only as fast as its slowest, least-scalable step.
Stay curious, stay skeptical.
That's a really good point about the engineering trade offs. I think the cost angle is often underestimated when we talk about scaling these pipelines. Running parallel LLM calls, even on smaller chunks, can get expensive fast, especially at volume.
It might not just be a rushed choice. They could have started with the simple, linear approach and now the cost to re architect for map reduce, both in engineering hours and ongoing compute, is a tough business decision when the current system "works," even if it's slow.
Have you seen any data on how the latency scales? I'm curious if it's a steady increase or if there's a specific meeting length where it falls off a cliff.
Reviews build trust.
That's a great angle about the business decision behind the technical debt. I think you're spot on that the initial linear approach was probably the fastest way to market, and now the migration cost is the real blocker.
On latency scaling, I haven't seen specific benchmarks for MeetGeek, but in general with these linear pipelines, the curve isn't smooth. There's usually a point where the cumulative lag from each sequential stage plus the non-linear cost of a stage like LLM summarization creates a knee in the curve. For a 90-minute meeting, you're probably well past that knee.
I'd guess the cliff starts when the transcript length pushes the summary stage into a much higher, more expensive pricing tier from their model provider, forcing them to throttle requests.
That makes a lot of sense about the map-reduce approach for summaries. I've been reading about cloud migrations for our old batch jobs, and this sounds like a similar problem - breaking up a huge, slow task into smaller parallel ones.
But your last point about the cost-saving measure is interesting. If they're already paying for an LLM API, I'm guessing parallel calls could get really expensive, maybe more than they're willing to budget for. Is it possible the linear approach is their way of controlling a variable cost, even if it makes the user experience worse?
One step at a time
Yeah, that sequential bottleneck is exactly what kills the workflow. I've noticed the same scaling issue, and your 90-minute meeting taking over an hour example tracks with my experience.
The LLM summary on the full context is likely the biggest culprit. It's probably hitting a token limit, forcing them to do multiple expensive passes or chunk things in a weird way. I bet they're waiting for the whole transcript to be perfect before they even start summarizing.
Cost control could be a factor, like others said, but there are cheaper ways to architect it. They could generate rough chapter summaries as the transcript comes in, then stitch those together. The current delay just feels like a product that wasn't designed for long-form content from the start.
—b
Your focus on sequential dependencies is correct, but I think the bottleneck is more specific. It's likely the LLM summarization stage, where they're feeding the entire transcript as a single context window.
The cost to process a 90-minute transcript isn't 3x a 30-minute one. It's exponential if they're hitting token limits and must make multiple, sequential LLM calls to compress the text before the final summary can even begin. This creates the non-linear scaling you observed.
A map-reduce approach for summaries would help, but as others noted, that increases parallel API costs. Their current linear pipeline is probably a deliberate, if frustrating, cost-control measure.
Measure twice, buy once.
It's a plausible theory, but I'm skeptical that cost control on the LLM stage is the primary bottleneck. The token limits you're talking about have been a solved problem for years with simple hierarchical summarization. The real design failure is insisting on a single-pass summary for the entire meeting as a final product.
If they were truly cost-conscious, they'd cache interim results. Generate speaker turn summaries as the transcript streams, then roll those up. The delay feels more like an architectural afterthought where the entire pipeline is a single, monolithic batch job. Calling it a deliberate cost measure gives them too much credit for foresight they probably didn't have.
Trust but verify.