A meeting, an interview or a study group produces a transcript where everything runs together. Speaker identification listens to the recording a second time, works out how many distinct voices there are and where each one starts and stops, then lines that up with the words so every turn is labelled.
Running it
Transcribe the recording as usual, then press Identify speakers in the footer of the transcription dialog, above Insert on mobile. It asks how many people were talking. If you know, say so; it is the single thing that helps the most, because otherwise the model has to guess the count from how alike the voices sound, and two similar voices can easily come out as one or three. Detect automatically is there for when you do not know.
A Speakers tab opens beside the Original one and fills in with the result, with the number of voices it found in the tab name. Each change of voice becomes its own paragraph, starting with Speaker 1, Speaker 2 and so on in the order they were first heard. If the count looks wrong, close the transcription, open it again and run it with a different number.
The second listen takes about a sixth of the recording's length, so a ten minute call is ready in under two minutes and an hour long meeting in about ten. The Speakers tab shows the wait, and on a phone it carries on from where it was if the screen locks in between.
The Speakers tab behaves like any other: edit it, run a quick prompt against it, and Insert puts the labelled version into the note. The Original tab stays as it was, so nothing is lost if the split turns out not to be what you wanted.
Naming the speakers
Grape does not know who anyone is; it only tells voices apart, and Speaker 1 is whoever spoke first. The Speakers tab shows one name field per voice it found. Type a name and every turn by that speaker is relabelled as you type, so John: and Jack: replace Speaker 1: and Speaker 2: before anything reaches the note. Clear a field and the number comes back. Names typed after you have edited the text by hand are applied to the labels only, so your edits stay.
What it is good at
Two to five people taking turns, each with a reasonable microphone: a recorded call, a one on one, a seminar with questions. Speakers are assigned a sentence at a time, so a change of voice is placed where the sentence ends rather than mid thought. Crosstalk, people speaking over each other and a single microphone far from the group all make the boundaries fuzzier, and someone who finishes another person's sentence will be filed under the person who started it.
Plans and models
Speaker identification is part of Grape AI on the Pro plan. The voice separation runs on the Grape server, so it works the same whichever transcription model you have enabled, including one from your own provider, and it always uses Grape's own transcription for the pass that needs word timings. It counts towards the transcription allowance like a normal transcription does.
Related
- Transcribe a recording or audio file
- Record system audio on desktop for calls that play through your computer