Whisper large-v3 with clean audio is close to a human transcriber; with poor audio it makes mistakes like anyone would. The six rules below matter more than any setting.
Transcribe now, free →One voice at a time: when two people overlap, no engine transcribes well. Audio that is not over-compressed: original m4a or wav beats a 32 kbps mp3 re-encoded several times. Normal volume: if you have to max the volume to hear, the engine hears as badly as you do.
Converting the file to other formats, "cleaning" the audio with aggressive filters that remove the voice too, re-uploading the same file several times. If the text still comes out poorly, the problem is almost always the recording: replay ten seconds and ask whether a stranger would understand the words.
With clean audio in one of the 8 supported languages, usually fewer than one word in twenty needs fixing; with noise or overlapping voices, much less.
Only slightly and only in the first seconds: if you know the language, choosing it is safer.
Yes: copy the text and fix names and numbers using the timestamps to find the spots in the audio.