Automatic subtitles
Speech recognised and timed into a subtitle file, on your own computer.
Drop files here
or click to choose from your deviceAccepted: MP3, M4A, AAC, WAV, OGG, OGA, OPUS, FLAC, WEBA, MP4, M4V, WEBM, MOV, MKV
Up to 3 files per runUp to 20 MB per file10 runs per dayRemove limits with Pro✓ Works offline
Your files stay on your device
This tool runs entirely in your browser. Nothing is uploaded, nothing is stored, and you can even turn off your internet connection before you start.
How it works
- Add the video or the recording. The sound is read; the picture is left alone.
- Choose the language spoken — nine are offered — and the subtitle format you need.
- Tick “Separate the speakers” if several people talk; each subtitle then begins with the speaker it belongs to.
- Wait while the recognition model downloads once and then runs. A long video takes a while.
- Download the SRT and load it into your editor, your player, or the “add subtitles to a video” tool to burn it in.
Frequently asked questions
- Where do the timings come from?
- From the speech itself: the recording is divided at its pauses, each stretch is recognised, and the subtitle takes the start and end of the stretch it came from. That is why a line appears when it is spoken rather than on a fixed cadence.
- SRT or WebVTT?
- SRT for video editors, players and most upload forms. WebVTT is what a browser’s own subtitle track reads, so use it for a video on a web page.
- How good is the recognition?
- Good on clear speech from one speaker. Crosstalk, heavy background noise and proper nouns are where it slips, so read the file through before you publish it — it is a plain text file you can edit in anything.
- Can the subtitles say who is speaking?
- Tick “Separate the speakers” and each subtitle begins with Speaker 1, Speaker 2 and so on. The recording is then read a second time by two further models — one that finds where the voice changes, one that turns each stretch of speech into numbers standing for the voice — so the first run downloads another 34 MB and the wait roughly doubles. They are numbered rather than named because nothing in the recording says who anybody is; replacing “Speaker 1” with a name afterwards takes a moment. Two people recorded clearly come out right; one microphone in a meeting room, voices that sound alike and long stretches of people talking over each other are where it makes mistakes.
- Is my video uploaded?
- No. The model is downloaded to your browser and the recognition happens there. The video never leaves your computer.
Made for a specific use
Related tools
❝ Japanese speech to text Japanese recordings and videos into a transcript and SRT subtitles, without the audio leaving your device. ❝ Meeting recording to minutes Who said it, when they said it, and the headings ready for what you decide it meant. ▭ Add subtitles to video An SRT file burned into the picture, so the words show on every screen. ≋ Remove background noise from audio Fans, hiss and hum taken out from under a voice.