Where the work happens
Speaker identification runs on your own machine. The models it needs — a segmentation model and a voice-embedding model — download automatically the first time you use the feature, and everything after that is local. That matters for two reasons: it keeps working when the call audio never leaves your device, and it means the labels aren’t a cloud feature you’re waiting on. Where your data goes has the full picture.Turning it on or off for good
The master switch is Identify and label speakers, and it’s on by default. Open Settings, then Speech-to-Text under AI Models, and pick the Note Recording tab — the toggle sits underneath the four engine options. Switch it off and transcripts show “You” and “Others” instead of named speakers, as the setting itself says.It’s on the Note Recording tab specifically, alongside the engine used for
meetings — not in a section of its own. If you’re looking for a “Meetings”
section in Settings, there isn’t one.
Controls during a recording
You can also override the setting for a single meeting without changing it globally. The pill at the top of the transcript carries the controls:
The count is a hint, not a limit you have to get right. When the note came from a
calendar event, the attendee list pre-fills it. OpenWhispr supports up to 8
speakers in a recording and starts from an assumption of 2.
Getting the count roughly right genuinely helps. Setting it far too high is how
you end up with phantom speakers — one person split across three labels.
Naming people
Labels start as Speaker 1, Speaker 2 and so on. Click one and type a name to replace it. If the meeting came from a calendar event, its attendees are offered as suggestions, and people you’ve named before appear under Known speakers. Once you name someone, that voice can carry the same name into later meetings rather than starting from scratch each time. Labels also carry a state, which tells you how much to trust them:Fixing a wrong label
Select the mislabelled segments and assign the right person. You can select several at once — the toolbar shows “3 selected” with an Assign to… action, which is much faster than correcting line by line. Corrections you make are treated as confirmed, so post-processing won’t undo them.If the labels look wrong mid-call
Give it until the end. Live labelling works from a fraction of the audio; the refinement pass at the end of the call sees the whole recording and regroups the voices properly. The app says so itself while recording: “Still identifying speakers. Labeling accuracy will improve once the call ends.” Things that genuinely hurt accuracy, in rough order:- Everyone on one microphone in a room. Voices sharing a channel are the hardest case there is.
- A speaker count set far from reality.
- Heavy crosstalk. People talking over each other blurs the boundaries.
- Only capturing your own microphone. If the other side was never recorded, there’s nothing to label — see capturing both sides of the call.
Turning it off
Two levels, depending on what you want:
Either way the transcript still captures everything — it just labels by source,
“You” for your microphone and “Them” for the call, instead of by voice.
There’s no separate cost to leaving speaker labels on, and no plan tier to
reach first — the models run locally on your machine.