Any audio in a note can become text you can search, quote and edit. Transcription is part of the Pro and Lifetime plans, using Grape AI or a transcription model from your own provider. On the free plan, Get started with Grape walks through the upgrade and a first transcript.
Transcribing
Open the menu on a recording or audio card and choose Transcribe. Grape sends the audio to your transcription model and opens the result in a dialog, on the Original tab, where you can edit it before doing anything else.
You can close the dialog while it works. The transcription carries on, the recording in the note says it is transcribing, and Grape tells you when the text is ready, see Get notified when the AI is done. The finished text is kept on the device that made it, ready to insert: open the recording again, or on the phone tap Transcript ready on it or choose Open Transcript from its menu, and the text is there with nothing sent a second time. A phone keeps the text for thirty days.
A voice note recorded from the iPhone Lock Screen needs none of this on Pro and Lifetime: Grape transcribes it on its own after you stop and puts the text in the new note under the recording.
Long recordings
A lecture or a meeting of an hour or more transcribes like a short memo. Anything longer than four minutes is sent in parts of up to four minutes, three at a time, and the text comes back as one transcript. Each part starts a few seconds before the previous one ends, so a word cut in half at a join is heard whole on one side of it and nothing goes missing.
The dialog counts the parts as they come back, and on desktop so does the recording in the note. To stop a long one early, choose Cancel in the dialog on desktop or Stop transcribing on the phone, and the parts still to go are never sent.
On Grape AI a single recording can run to about six and a half hours, which is as much as one five hour window of the transcription allowance holds, and a month holds about 33 hours. That is for recordings made in the app and compressed files like MP3 or M4A: an uncompressed stereo WAV goes in parts of two minutes or less to stay under the upload limit, so it uses the allowance at least twice as fast. With a transcription model from your own provider, the only limits are the ones that provider sets.
Quick prompts on the transcript
If you have prompt templates enabled, they appear as chips in the transcription dialog. Tapping one, like Summarize, runs it against the transcript and opens the result in its own tab, so the raw transcript and the summary live side by side. Insert puts whichever tab you are viewing into the note.
Who said what
A recording with more than one voice can be split by speaker. The Identify speakers button in the same dialog opens a Speakers tab with one paragraph per turn, labelled Speaker 1, Speaker 2 and so on. See Identify speakers in a recording.
Dictation on desktop
With a transcription model from your own provider enabled, a microphone button appears in the desktop editor header. It listens and types into the note as you speak, which suits thinking out loud better than typing does.
Accuracy
Transcription is good enough to search and work with, and not good enough to publish unchecked. Names, technical terms and crosstalk are the usual weak points, so listen back before you rely on a specific line.