Home · Guides

How to improve the accuracy of an automatic transcription

Whisper large-v3 with clean audio is close to a human transcriber; with poor audio it makes mistakes like anyone would. The six rules below matter more than any setting.

Transcribe now, free →

How it works, step by step

  1. Record well: microphone within a metre, quiet room, no speakerphone in large rooms. 90% of errors start here.
  2. Choose the language instead of auto-detect when you know it: it removes the rare cases where the first chunk is misread.
  3. Proofread proper names, numbers, acronyms and technical terms: these are the words the engine cannot know. The timestamps take you straight to the spot to check.

What really helps

One voice at a time: when two people overlap, no engine transcribes well. Audio that is not over-compressed: original m4a or wav beats a 32 kbps mp3 re-encoded several times. Normal volume: if you have to max the volume to hear, the engine hears as badly as you do.

What does not help

Converting the file to other formats, "cleaning" the audio with aggressive filters that remove the voice too, re-uploading the same file several times. If the text still comes out poorly, the problem is almost always the recording: replay ten seconds and ask whether a stranger would understand the words.

FAQ

What accuracy can I expect?

With clean audio in one of the 8 supported languages, usually fewer than one word in twenty needs fixing; with noise or overlapping voices, much less.

Is auto-detect less accurate?

Only slightly and only in the first seconds: if you know the language, choosing it is safer.

Can I improve the result afterwards?

Yes: copy the text and fix names and numbers using the timestamps to find the spots in the audio.

Transcribe now, free

No account · no cost · in seconds

Transcribe now, free →

More guides