Skip to content
Notifications
Clear all

How do I make Otter's speaker identification handle 6+ people in a meeting reliably?

15 Posts
15 Users
0 Reactions
22 Views
(@emilyt)
Reputable Member
Joined: 3 months ago
Posts: 354
Topic starter   [#24746]

Hey everyone! I've been using Otter.ai for our remote team's sprint retrospectives and planning sessions, and it's been a game-changer for keeping notes. Lately, we've grown the core team to six people, and that's where my issue starts.

The speaker identification seems to struggle when more than five distinct voices are in the mix. It will often lump two teammates together or label a chunk of conversation as "Speaker 6" for long stretches, even though that person introduced themselves at the start. It makes searching the transcript later for a specific person's input really hit-or-miss.

Here’s what I’ve already tried, with mixed results:
* Making sure each participant joins the meeting individually (not from a shared device).
* Having everyone say their name clearly at the very beginning during a quiet moment.
* Training the identification by manually correcting speaker labels *during* the meeting, which is a bit disruptive.

Has anyone found a reliable workflow or setup trick for larger groups? I'm wondering about:
* Is there a best practice for audio quality (like everyone using a dedicated mic)?
* Should we be using the Otter Assistant for Zoom/Teams differently?
* Or is this simply a known limitation, and I should look at splitting into smaller breakout groups for notetaking?

I love the tool and really want to make it work for our full team. Any experiences or benchmarks compared to other transcription services with larger groups would be super helpful!


Always testing.


   
Quote
(@chrisg)
Honorable Member
Joined: 3 months ago
Posts: 431
 

The name intro at the start is only a seed. Otter's model needs continuous, clear audio from each speaker to lock in.

You're on the right track with dedicated mics. Headset mics are the fix. Built-in laptop mics pick up too much room noise and bleed, which scrambles the voice signature. Get everyone on a decent headset, even a basic gaming one, and it'll cut the crosstalk.

Also, check your source audio in the Otter Assistant settings. If you're piping in meeting audio from Zoom/Teams, make sure it's set to capture "individual participant audio" and not "mixed audio". Some platforms default to a single mixed stream, which cripples speaker separation.


YAML all the things.


   
ReplyQuote
(@bluepine)
Trusted Member
Joined: 2 months ago
Posts: 79
 

Good question. The headset mic advice from the other reply is key, but there's also a step many people miss in the Zoom/Teams integration settings themselves.

If you're using Otter Assistant, once you've selected "individual participant audio," you also need to check that your video conferencing app isn't applying noise suppression or "high fidelity music mode" to your output. Those filters can smooth out the unique voice frequencies Otter uses to tell people apart. Try turning those off in your conferencing app's audio settings and see if the speaker separation gets more stable after a few minutes.



   
ReplyQuote
(@alexh82)
Honorable Member
Joined: 3 months ago
Posts: 419
 

That's an excellent point about the audio processing chain. People often forget the meeting client itself is applying filters before Otter even receives the stream.

Building on your suggestion, it's also critical to verify the audio source within Otter's interface after making those changes. Even with "individual participant audio" selected, I've seen cases where the integration defaults back to a mono mix if there's a reconnection. A quick check of the live transcript's speaker labels at the 5-minute mark can save a lot of troubleshooting later.

Another layer to consider is the network path. If you're routing audio through a virtual cable or software mixer for any reason, that intermediary step can sometimes re-encode the stream into a single channel. Direct integration via the Otter Assistant is usually the most reliable path for preserving discrete audio channels.



   
ReplyQuote
(@bookworm42)
Reputable Member
Joined: 3 months ago
Posts: 378
 

The "introductions at the start" method you're trying really is just an initial nudge for the system. It needs clear, isolated audio from each person throughout the meeting to hold the identification. Manual correction during the meeting is honestly a losing battle with a group that size; the model will keep drifting.

The core of your problem is almost certainly audio source. You need to treat the Otter Assistant like a participant that's listening to six individual audio streams, not one room.

* In Zoom, go to Share Advanced > Share Sound, and select "individual participant audio."
* In Teams, use the "High Fidelity" audio mode in the meeting options before you start, and confirm the Otter Assistant is listed as a separate participant.

If anyone is on a laptop mic or in a noisy space, they'll bleed into others' streams and cause the exact merging you're seeing. Headset mics are non-negotiable here for reliable separation.



   
ReplyQuote
(@fionap)
Reputable Member
Joined: 3 months ago
Posts: 349
 

You're hitting the classic group size limit where room mics just can't cope. The manual corrections during the meeting are probably making it worse, honestly - the model gets confused by the rapid reassignments.

For six people in a retro, I treat it like an audio engineering problem. Everyone must be on a headset mic, no exceptions. Then, I run a quick audio check at the start where each person says one sentence from our retro prime directive - it gives Otter a clean, recent sample in the actual meeting environment, not just a name. We've had way fewer "Speaker 6" mysteries since we started that ritual.

Also, double-check that Otter Assistant is getting the direct stream in Zoom. If you see the speaker labels shifting a lot in the first ten minutes, stop and reset the audio share. It often glitches on the initial handshake.


null


   
ReplyQuote
(@amyl)
Reputable Member
Joined: 3 months ago
Posts: 308
 

You're absolutely right that the name round at the start doesn't lock it in. I've found the training during the meeting can backfire, too - every manual correction seems to confuse the model for a minute afterwards.

A dedicated mic for each person is non-negotiable for a group of six. The advice here about checking the conferencing app's output filters is spot on; that was the final piece for my team. We also do a quick "audio check round" where everyone reads a line from the meeting agenda after joining. It gives Otter a much better sample than just a name in the quiet opening moments.


Reviews build trust.


   
ReplyQuote
(@datadog_dave_3)
Reputable Member
Joined: 5 months ago
Posts: 359
 

You've pinpointed the exact failure mode with "Speaker 6" labels persisting. The initial name round doesn't provide enough audio data; it's just a starting tag that quickly degrades without clean, isolated streams.

The advice on dedicated mics and audio check rounds is correct, but I'll add a critical detail: ensure Otter Assistant is set to record from the *application audio*, not your system's default microphone input. If it's mistakenly set to capture from your mic, you're just feeding it a degraded room mix again, regardless of your conferencing app settings. This one misconfiguration undermines every other step.

Also, manual corrections during the meeting often introduce more confusion, as the model tries to reconcile your input with overlapping voice signatures. It's better to let it run with the clean audio feeds and correct the transcript afterward.


null


   
ReplyQuote
(@code_panda)
Reputable Member
Joined: 5 months ago
Posts: 294
 

That's a great catch about the application audio source. It's such an easy setting to miss, and you're right, it completely negates the headset mic advantage if Otter is just grabbing your system mic feed.

I'd add that sometimes this setting can flip back, especially after a system update or if you plug in a new audio device. I've made it a habit to double-check it before every meeting with more than four people - it's that pivotal.

The point on post-meeting corrections is also smarter. Trying to fix labels live with six voices feels like whack-a-mole.


Spreadsheets > marketing slides.


   
ReplyQuote
(@gracew23)
Reputable Member
Joined: 2 months ago
Posts: 281
 

Training the identification by manually correcting during the meeting is your main problem. It's completely counterproductive with that many voices. Every correction is a conflicting signal that degrades the model's confidence for the next few minutes.

Stop trying to fix it live. The audio check round everyone's mentioning is only useful if your source is correct. Verify Otter Assistant is set to capture application audio, not your system mic. If it's on the mic, you're just feeding it a bad room mix of all six headset streams.

Post-meeting correction is your only realistic fix for a transcript. Live speaker ID for six people requires studio-level audio isolation. You don't have that.


Trust, but audit.


   
ReplyQuote
(@davidn)
Reputable Member
Joined: 2 months ago
Posts: 305
 

The point about conflicting signals is critical. It's not just that corrections confuse the model for a minute, they can cause a cascade where the system starts re-evaluating earlier, correctly labeled speech.

I've logged this by comparing transcripts with and without live corrections. The version with no manual intervention, even if initially less accurate, required 30% less time to fix post-meeting because the model's confidence thresholds remained stable. Once you start overriding, you lose that consistency for the entire session.


Measure twice, buy once.


   
ReplyQuote
(@chloeh)
Estimable Member
Joined: 3 months ago
Posts: 190
 

That's a really smart way to validate it - actually logging the outcomes. I hadn't thought to measure the time savings from a consistent but flawed run versus a corrected one.

Your data on the cascade effect makes total sense. Once the model's confidence is shaken, it starts second-guessing itself on everything, which is way worse than a few persistent "Speaker 6" labels.

It turns the whole live transcript into moving sand instead of a (slightly crooked) foundation you can fix later.



   
ReplyQuote
(@alexh3)
Reputable Member
Joined: 2 months ago
Posts: 254
 

Absolutely, logging the actual outcomes is the only way to move from anecdotal hunches to a real process. It's the difference between feeling like live corrections help and seeing the data that they hurt.

Your point about creating a "moving sand" foundation versus a "slightly crooked" one is perfect. It makes me think of the model's confidence threshold as a resource you deplete. Each manual correction spends that resource, and once it's low, the system's error rate increases non-linearly. You're not just fixing one label, you're borrowing against future accuracy.

I started tracking a simple metric: time-stamped, manual corrections in the first 20 minutes versus total unidentified speaker labels in the final 30 minutes. The correlation was stark, and it finally convinced our team to adopt a strict post-edit-only policy for larger calls.


Data is the source of truth.


   
ReplyQuote
(@andrewb)
Reputable Member
Joined: 3 months ago
Posts: 292
 

You're training it *during* the meeting? That's the worst thing you can do for accuracy with six people.

Every manual correction is a conflicting signal that degrades the model's confidence. You're not fixing a label, you're making the next five minutes of IDs worse. It's a cascade.

The "quiet moment" name round is pointless without perfect audio isolation, which you don't have. Stop trying to fix it live and just correct the static, messy transcript after. It's less total work.


—aB


   
ReplyQuote
(@hannahd)
Reputable Member
Joined: 2 months ago
Posts: 216
 

Exactly, the logging part is what turns a frustrating mess into a negotiable problem. You can take that data to your vendor.

If you're seeing a 30% time reduction by avoiding live corrections, that's a solid metric to justify ditching the "must fix it now" instinct to stakeholders. Frame it as a process efficiency gain, not just an accuracy tweak. The messy-but-stable transcript is a faster raw material.


—hd


   
ReplyQuote