ResourcesFile Transcription

How to Batch Transcribe Audio Files in Different Languages on Mac Offline

Prepare local models, choose automatic detection or language-based groups, then review and export each recording in a multilingual queue.

2026-10-04 · 7 min
A queue of colourful audio cards leads to separate transcript documents on a warm ivory background.

The short answer

To transcribe several recordings in different languages on a Mac, prepare the local recognition resources, add the files to a queue, and decide how to handle their spoken languages. Automatic language detection is useful when a compatible model is installed. If you know the language and detection is unreliable, group similar files and run them with an explicit source-language setting.

Alison Workspace supports file queues, multilingual recognition choices, optional translation, and speaker separation. It requires an Apple-silicon Mac, M1 or newer, with macOS 15 or later. After the required resources are ready, supported file recognition runs locally. Plan around the fact that the queue shares settings for a run: several files in different languages are not the same problem as one recording that switches language repeatedly.

Organize the files around the job you need done

Suppose you have an English lecture, a Japanese interview, and a Spanish voice recording. You want a separate original-language document for each, with a translation only where it helps. This is an illustrative workflow, not a report of a three-language accuracy test.

Keep the files named so that you can tell them apart without opening every transcript. Check their sound, approximate length, and main language. For audio embedded in a video, the file still needs an accessible, decodable audio track; its picture does not supply words to speech recognition.

Decide whether you need source text, a translated reading aid, speaker labels, or subtitles. Each option adds a separate requirement and a separate review task. Starting with original text can make it easier to identify a language problem before adding translation.

Prepare recognition resources before going offline

The recognition menu offers many languages, but the engine and model must support the language you intend to use. A language appearing in a menu does not guarantee that every installed model can handle it equally well. Use a multilingual model for a batch that needs multiple source languages.

Alison's automatic file-language detection requires a usable local Whisper model. Some explicitly selected languages can use available system recognition on compatible macOS versions, but that is not an automatic-language substitute. Follow the app's readiness message instead of assuming that a model used by another workflow is ready here.

Download and prepare resources while connected, then try a short file with the intended settings before relying on an offline session. First preparation can take time. No model-ready check can prove the quality of an unfamiliar accent or a noisy recording.

Choose automatic detection or language-based groups

For files that each have a clear main language, automatic detection can avoid manually sorting every item. It can still be wrong, particularly with silence, music, very little speech, or an unusual opening. Inspect the result for every file rather than accepting the whole queue because the first transcript looks right.

If you know the languages, a more controlled approach is to run the English files with English selected, then the Japanese files with Japanese selected, and so on. The queue uses a shared recognition choice for that run; this guide does not describe a separate language setting attached to every queued file.

Do not start a mixed-language batch with a fixed language unless you actually want that language for all files. When recognition looks wrong, check the audio and language first. Changing the target translation language cannot repair a source transcript produced under the wrong recognition setting.

Start a manageable batch

Begin with a small set so you can confirm the settings, resource readiness, and output before processing a larger collection. Multiple files are queued for processing; this is not a promise that all files run simultaneously or finish at a fixed speed.

  1. Choose the files you can read locally and confirm their names, durations, and expected spoken languages.
  2. Open the file-transcription window and add several files, or drag them into its queue.
  3. Select automatic language detection with a ready multilingual model, or group files by language and choose an explicit source language.
  4. Leave translation off for original-language documents, or choose one supported translation target for the run and prepare its resources.
  5. Enable speaker separation only when useful and prepare its resources; use an expected speaker count only if it fits the files in that run.
  6. Start the queue, then review each file's status and generated text before exporting its document.

Check any rejected file as a media issue first. A file with no audio track or unreadable encoding will not be fixed by adding more languages to the model settings.

Keep recognition, translation, and speaker settings distinct

Recognition produces words in the language spoken. Translation turns those words into another reading language. Speaker separation groups text by voice. They may be used together, but success in one does not establish success in the others.

For example, a correctly recognized Japanese passage may still have an unavailable translation target or a wrong speaker assignment. Keep the original visible when reviewing a translated document. Verify names, figures, and important negative statements against the recording, especially when the translation seems smoother than the source text.

Translation pairs depend on resources available on the actual Mac. A multilingual recognizer does not imply that every source language can be translated into every target. If translation is unavailable, you can still use and review a source-language transcript rather than mislabelling it as a translation.

Handle a recording that changes language differently

A folder of mostly single-language files is a good reason to consider automatic detection. One interview with repeated switches between Japanese and English needs separate attention. Automatic detection does not establish that every switch inside the recording was recognized correctly.

Review passages around the language changes. If the result consistently loses one language, consider processing suitable excerpts separately with the intended language selected, using your normal media-editing tools to prepare them. Preserve the original and remember that a transcript of an excerpt uses that excerpt's timeline.

There is no claim here of flawless mixed-language recognition or equal quality in every supported language. Background music, regional speech, unfamiliar names, and different recording conditions also affect what needs correction.

Check completion and retry only after diagnosing the issue

Alison's file-transcription trial currently provides 60 minutes of cumulative audio duration, separate from live listening. The queue uses the remaining allowance in order. Later files may receive only the remaining time or none; look at the expected and processed lengths rather than assuming that every added file will finish completely.

A partial result is not a complete transcript of a long recording. Failed or cancelled jobs can be retried after you check the cause. If a file was processed under the wrong language, select the correct language for another run; if resources are missing, prepare them first. Review the allowance before repeating a long batch.

If you unlock complete-file access, run the affected file again. Changing access does not automatically fill the unprocessed ending of an earlier document. Keep the relationship between the file, its status, and its latest transcript clear.

Review and export each document intentionally

Use search and playback to inspect important passages in each result. Editing one transcript does not mean the other documents in the queue have been reviewed. Give exported files distinct names and keep the original recordings available for later checking.

The current transcript export menu offers TXT, Markdown, CSV, PDF, SRT, and VTT. Choose original text, translation where available, or both; speaker and document-timestamp options depend on the record and format. Export each record through its menu rather than assuming that adding a batch provides a one-click bulk export.

Local recognition describes where the processing happens. It does not control what a synced export folder, shared drive, or backup later does with the files. Choose where you keep your recordings and documents. The next step is a small queue whose languages and outputs you can verify before you scale up.

Official references checked October 4, 2026

WhisperKit: on-device speech-to-text and model choices
Apple: file-based media reading
FluidAudio: local speaker diarization

Frequently asked questions

Can different-language recordings go into one queue?

Yes, but the run shares language settings. Use automatic detection with a ready compatible multilingual model, or group files by language and run those groups with an explicit source language.

Does automatic language detection need a model download?

It needs a usable local Whisper model. Prepare that model before an offline session. A system engine for one explicitly selected language is not an automatic-language fallback.

Will it reliably handle two languages switching inside one file?

Do not assume that. Review the switching passages and consider separately prepared excerpts with the correct language settings when one language is repeatedly missed.

Does transcribing in Japanese automatically give me English text?

No. Recognition and translation are separate choices. Select English as a supported translation target and prepare the required pair if you want an English reading aid.

Can I transcribe files without internet after preparation?

Supported recognition runs locally once its required resources are usable. Prepare any optional translation and speaker resources too, and confirm the intended workflow before going offline.

Are speaker labels available for a batch?

Speaker separation can be enabled for the run when its resources are ready. Review each file independently; speaker numbers do not establish the same person's identity across files.

Does batch transcription include automatic bulk export?

The queue processes several files. Export is available from each transcript's menu; this guide does not claim a one-click export of the whole queue.

Related guides