Nouplo

Japanese speech to text

Japanese recordings and videos into a transcript and SRT subtitles, without the audio leaving your device.

Drop files here

or click to choose from your deviceAccepted: MP3, M4A, AAC, WAV, OGG, OGA, OPUS, FLAC, WEBA, MP4, M4V, WEBM, MOV, MKV

Up to 3 files per runUp to 20 MB per file10 runs per dayRemove limits with Pro✓ Works offline

Your files stay on your device

This tool runs entirely in your browser. Nothing is uploaded, nothing is stored, and you can even turn off your internet connection before you start.

How it works

  1. Add an audio or video file: MP3, M4A, WAV, MP4 and most others a browser can play.
  2. Tick “Put the time at the start of each line” if you want timestamps in the transcript.
  3. Tick “Separate the speakers” if more than one person is talking; each line then begins with Speaker 1, Speaker 2 and so on.
  4. Select Transcribe. The first time, the speech model (about 160 MB) downloads once; after that it starts straight away.
  5. Copy the transcript, or download it as a text file together with SRT subtitles.

Frequently asked questions

How accurate is it?
It uses ReazonSpeech (Reazon Research, Apache 2.0), a model trained on 35,000 hours of Japanese. In our tests it got about 6 % of characters wrong on read speech and about 4 % on television audio, counting spelling variants such as 私 against わたし as errors. Clear speech comes out best; noise, people talking over each other, jargon and rare names bring more mistakes.
Can it transcribe English?
Yes: switch the language to English, or use Transcribe audio to text. English is read by a different model, Whisper base, because each language gets the model that measured best: for Japanese that is ReazonSpeech, which made about half the errors of the largest Whisper a web page can carry, at least eight times faster.
Is there punctuation?
No. The model writes none, so the transcript starts a new line wherever the speaker paused, which keeps it readable; the subtitles are cut at the same places.
Can it tell who is speaking?
Tick “Separate the speakers” and every line begins with Speaker 1, Speaker 2 and so on. The recording is then read a second time by two further models — one that finds where the voice changes, one that turns each stretch of speech into numbers standing for the voice — so the first run downloads another 34 MB and the wait roughly doubles. They are numbered rather than named because nothing in the recording says who anybody is; replacing “Speaker 1” with a name afterwards takes a moment. Two people recorded clearly come out right; one microphone in a meeting room, voices that sound alike and long stretches of people talking over each other are where it makes mistakes.
Is my recording uploaded?
No. Recognition runs in this page, on your own device, which makes it suitable for meetings, interviews and other recordings that should not leave it. Once the model has downloaded, it also works offline.
How long does it take?
On a computer, roughly a tenth of the recording's length: a few minutes for an hour. Phones are slower, and a long recording can run out of memory there; cut it into shorter parts first.

Made for a specific use

Related tools