Nouplo

Transcribe audio to text

English recordings into a punctuated transcript and SRT subtitles, on your own device.

Drop files here

or click to choose from your deviceAccepted: MP3, M4A, AAC, WAV, OGG, OGA, OPUS, FLAC, WEBA, MP4, M4V, WEBM, MOV, MKV

Up to 3 files per runUp to 20 MB per file10 runs per dayRemove limits with Pro✓ Works offline

Your files stay on your device

This tool runs entirely in your browser. Nothing is uploaded, nothing is stored, and you can even turn off your internet connection before you start.

How it works

  1. Add an audio or video file: MP3, M4A, WAV, MP4 and most others a browser can play.
  2. Tick “Put the time at the start of each line” if you want timestamps in the transcript.
  3. Tick “Separate the speakers” for an interview or a meeting; each line then begins with Speaker 1, Speaker 2 and so on.
  4. Select Transcribe. The first time, the speech model (about 78 MB) downloads once; after that it starts straight away.
  5. Copy the transcript, or download it as a text file together with SRT subtitles.

Frequently asked questions

How accurate is it?
It uses Whisper base (OpenAI, Apache 2.0). On clearly read English (the LibriSpeech test set) it got about 5 % of words wrong in our tests. Accents, noise, people talking over each other and specialist words bring more mistakes; for noisy recordings, try Remove background noise first.
Does it add punctuation?
Yes. Whisper writes capital letters, commas and full stops itself. Each line of the transcript is a stretch of speech between pauses, and the subtitles are cut at the same places.
Can it tell the speakers apart?
Tick “Separate the speakers” and every line begins with Speaker 1, Speaker 2 and so on, which is what makes an interview or a meeting readable. The recording is then read a second time by two further models — one that finds where the voice changes, one that turns each stretch of speech into numbers standing for the voice — so the first run downloads another 34 MB and the wait roughly doubles. They are numbered rather than named because nothing in the recording says who anybody is; replacing “Speaker 1” with a name afterwards takes a moment. Two people recorded clearly come out right; one microphone in a meeting room, voices that sound alike and long stretches of people talking over each other are where it makes mistakes.
Is my recording uploaded?
No. Recognition runs in this page, on your own device, which makes it suitable for meetings, interviews and other recordings that should not leave it. Once the model has downloaded, it also works offline.
How long does it take?
On a computer, roughly a quarter to a half of the recording’s length. Phones are slower, and a long recording can run out of memory there; cut it into shorter parts first.
Can it do other languages?
Eight more — switch the language. Japanese is read by a model made for it; German, French, Spanish, Portuguese, Italian, Korean and Turkish by Whisper small, because base, which is good enough for English, made up to twice as many errors on them in our tests. That larger model is about 250 MB the first time.

Made for a specific use

Related tools