How to Transcribe an Interview with Speaker Labels on Mac
Add speaker labels to an interview, check who said each passage, correct assignments and export a reviewed conversation on your Mac.
2026-10-04 · 7 min
The short answer
To transcribe an interview and distinguish the people speaking, use file transcription with speaker separation enabled. It adds speaker labels to portions of the recognized text, helping you review a question and its answer without treating the whole recording as one voice. Prepare the recognition and speaker resources before starting, then listen back to check the assignments.
Alison Workspace supports this workflow for readable audio and video files on an Apple-silicon Mac, M1 or newer, running macOS 15 or later. You can select an expected speaker count, review the generated transcript, change a paragraph's speaker assignment, rename labels, and export the text. Start with a clear exchange between the interviewer and guest rather than judging the result from music or several people talking together.
Know what speaker labels tell you
Speech recognition asks what was said. Speaker diarization asks which voice spoke during a portion of the recording. A transcript can have correct words with the wrong speaker, or a correct speaker label attached to an incorrectly recognized sentence. Review both dimensions.
The labels initially describe voices within that recording. They do not automatically supply verified personal names or establish that a person in another file is the same individual. Voice characteristics help group speech; matching real identities is a separate task.
Speaker separation in this guide means labelled text. It does not mean exporting an isolated audio file for each participant or removing one person's voice from an overlapping exchange. Decide whether you need a readable conversation record or separate audio tracks before choosing a workflow.
Begin with an interview recording you can review
Use the best available original recording, with speech that is easy to hear. Avoid making a compressed copy simply to reduce its size if the original already imports. A quiet guest, background music, and voices arriving from different microphones can increase the work needed to check the result.
Listen to a short exchange first. Note how many people actually speak, including an introduction or a producer's question if present. Knowing that the main interview has two participants does not prove that every part of the file has only two voices.
If separate participant recordings are already available, keep them as references. The current workflow labels the imported file; it does not describe automatically merging separate participant tracks into a synchronized conversation. Make sure you are using recordings you may transcribe and retain.
Prepare speaker separation and the expected count
Speaker separation has resources in addition to speech recognition. In Alison's file-transcription window, enable the speaker option and follow its model-readiness prompt. Turning on a setting without usable resources is not proof that the final record will contain labels.
The count menu offers automatic estimation and known counts from 2 to 6. For a genuinely two-person interview, choosing 2 tells the process how many voices to group. If you are not sure, use automatic and inspect the result. The known-count menu is a setting range, not a measured limit on the number automatic mode can handle.
A count is guidance, not a guarantee. If a third voice appears, or if one person changes microphone or speaking style, an incorrect fixed count may group passages badly. Check the full recording rather than assuming that the opening establishes every speaker.
Run the interview through file transcription
Choose the recognition language for the actual speech. If you need a translation, select its target independently and confirm that the required pair is available. You can leave translation off while checking names and speaker assignments in the original language.
- Add the interview audio or video file to the transcription queue.
- Check the accepted file's duration and select the spoken language.
- Prepare the required recognition resources; automatic language detection needs a compatible local model.
- Enable speaker separation and prepare the requested speaker resources.
- Choose a known speaker count when justified, or leave it on automatic, then start transcription.
- Open the result and replay passages from each participant before accepting the labels.
Let the file's status tell you whether processing completed. If the result is partial, review only the processed portion and deal with the missing remainder before presenting it as a complete interview.
Correct assignments before adding names
Locate a clear sentence from each voice and play its audio. Establish which initial label corresponds to the interviewer and which corresponds to the guest. Check later passages too: a correct opening label does not prove that all changes of speaker were handled correctly.
When a paragraph is assigned incorrectly, enter the transcript's editing mode and choose the correct existing speaker for that paragraph. Correct the words separately if necessary. A name change applies to the label throughout the record; it is not the right tool for moving only one mislabelled paragraph.
Very short replies and interruptions deserve attention. A recognized segment can span a speaker change, and the available timing information affects how precisely it can be divided. Do not assume every brief “yes” or overlapped remark will receive its own perfect label.
Rename labels once the voices are confirmed
Replace generic labels with confirmed names or roles such as Interviewer and Guest. Roles are useful when you do not need personal names. Keep spelling consistent with the interview's accompanying material rather than guessing from an automatically recognized introduction.
The speaker-name controls update the labels used by the record and its exports. Replay a named passage to confirm that a global name change did not attach the guest's name to the interviewer. When one real voice has been split into several labels, check their passages before assigning the same name.
Renaming a label records your editorial decision. It is not voice authentication and does not create an automatic identity profile for future files. Review a new recording's labels independently.
Export a conversation that remains easy to check
Choose TXT, Markdown, PDF, or CSV for a readable or reusable document. Where speaker labels exist, enable the option to include speakers; keep timestamps if the reader will need to return to the recording. Original, translated, and bilingual export choices depend on the text available in the record.
SRT and VTT are also available for a subtitle workflow and always include their required time axis. Check the intended display before sharing a subtitle file: a useful interview document and a comfortable video subtitle track need different formatting decisions.
Listen again to names, quotations, numbers, commitments, and uncertain attributions before sharing. A speaker-labelled draft helps you find these passages; it does not make an unchecked quotation reliable.
Troubleshoot labels rather than forcing a perfect story
If no labels appear, check that speaker separation was enabled and its resources were ready. Recognition can produce text while the optional speaker stage provides no usable assignments. A successful text result and successful speaker labelling are not the same status.
If labels are wrong, revisit the count setting and the audio. Similar voices, distant speech, changes in microphone sound, background noise, and overlapping talk can make grouping difficult. Use clear passages as anchors, correct what you can verify, and leave uncertain attribution visible rather than inventing certainty.
After resources are ready, the supported recognition and speaker work runs locally on your Mac. Preserve the source file for review and choose an intentional destination for the exported transcript. The next useful step is one clear interview exchange with labels you have personally checked.
Official references checked October 4, 2026
Microsoft: diarization and generic speaker identifiers
Amazon Transcribe: speaker labels and timestamps
FluidAudio: on-device speaker diarization framework
Frequently asked questions
Should I choose automatic or 2 speakers for a two-person interview?
Choose 2 when exactly two people speak in the file. If introductions or interruptions add other voices, use a count that reflects the recording or try automatic and review the assignments.
Does speaker separation automatically identify people's names?
No. It assigns voice labels within the recording. Confirm the people or roles yourself, then rename the labels in the transcript.
Can I correct who said one paragraph?
Yes. The transcript editor lets you change a paragraph's assignment to an existing speaker. Renaming a speaker changes that label throughout the record instead.
Why did one person receive two labels?
Voice grouping can be affected by short speech, noise, different microphones, or changes in delivery. Replay the relevant passages before deciding whether the labels belong to the same person.
Can I export each person's isolated voice?
This feature labels transcript text. It does not provide a separate audio-track export for each speaker or remove overlapping voices.
Will speaker names appear in the exported document?
Enable Include Speakers when the record has labels. Confirmed names are used in the export. Check the chosen document or subtitle format before sharing it.
Will the same speaker number identify the same person in another file?
No. Treat the labels in each recording independently. The workflow described here does not establish a cross-file voice identity.