Browsers block Web Workers (needed for model loading/transcription) when a page is opened via file://. Serve this folder over local HTTP instead, e.g.:
npx serve .
# or
python -m http.server 8000
Then open http://localhost:<port> in your browser.
Local Whisper Subtitles
Add subtitles to an MP4 entirely in your browser — transcribe with Whisper, edit the text, translate it,
then download SRT/VTT or burn the subtitles into the video. No upload, no server, no sign-up.
Initializing…
For a non-English language its recommended you set the language yourself.
Whisper model — nothing is downloaded until you pick one
The model you choose is downloaded once and then cached by your browser, so picking it again later is instant.
Switching to a different size downloads that one the first time you use it.
Generate subtitles for your MP4
Drag & drop an MP4 file here
or click to browse
Re-encodes the video with subtitles permanently drawn into the frames (hardsub). Runs locally via ffmpeg.wasm; can take a while for longer videos.
Edits update the video preview automatically. Don't change the timestamp lines or blank lines between entries — just the subtitle text.
Translates the text currently in the box above (including your edits), replacing it in place. Requires "Spoken language" above to be set to a specific language, not Auto-detect.
What this free MP4 subtitle generator does
Local Whisper Subtitles is a free subtitle generator that runs completely on your own device.
It uses OpenAI’s Whisper speech-recognition model, compiled to WebAssembly and
executed through Transformers.js,
to turn the speech in a video into timed subtitles. Burning subtitles into the picture is handled by
ffmpeg.wasm.
Because both run in the browser, your video is never uploaded, there is no account or API key, and there are no
file-size or minute quotas beyond what your own hardware can handle.
How to add subtitles to a video, step by step
Choose language and download a model. Leave the spoken language on Auto-detect or pick it explicitly, then click the model you want: Tiny for speed, Base for a balance, or Small for the best accuracy. Nothing is downloaded until you click.
Add your video. Drag an MP4 onto the drop zone or click to browse. The file is read locally.
Transcribe. The model you downloaded is cached by your browser and runs in a Web Worker, so the page stays responsive.
Edit and translate. Fix any misheard words straight in the subtitle box — the video preview updates as you type — and optionally translate the result to English or Swedish.
Export. Download a .srt or .vtt file, or burn the subtitles permanently into a new MP4.
Features
Automatic speech recognition for MP4 video using Whisper tiny, base or small
Fully private: 100% client-side, no upload, no server, no tracking
Auto-detect or choose English, Arabic, Spanish, French, German, Hindi, Chinese, Russian, Portuguese or Turkish
Editable subtitles with a live preview on the video
Subtitle translation to English or Swedish
SRT and WebVTT export
Hardcoded (burned-in) subtitles via ffmpeg.wasm
Completely free — no sign-up, no watermark, no limits
Frequently asked questions
Is my video uploaded anywhere?
No. Transcription, translation and re-encoding all happen inside your browser using WebAssembly and, where available, WebGPU. There is no backend to send a file to.
Which subtitle formats can I download?
SubRip (.srt) and WebVTT (.vtt). You can also burn the subtitles into the MP4 as hardsubs so they show in any player, including ones that ignore sidecar subtitle files.
Which languages are supported?
Whisper can auto-detect the spoken language, and English, Arabic, Spanish, French, German, Hindi, Chinese, Russian, Portuguese and Turkish can be selected explicitly. Finished subtitles can be translated to English or Swedish.
Which Whisper model should I use?
Tiny (~75MB) is the fastest and leans English-only. Base (~145MB) is a good balance and handles non-English audio well. Small (~480MB) is the most accurate but noticeably slower. You pick which one to download — the page never fetches a model on its own.
Is it really free?
Yes — completely free, with no account, API key or usage cap, because all the computation runs on your own machine.
Learn how any of this works
Three referenced explainers on the machinery behind this tool — speech recognition, machine translation
and the subtitle formats themselves. Every claim links to a primary source.
The six stages between a video file and a translated caption track — audio extraction, speech recognition, cue segmentation, translation, re-timing, rendering — and how to tell which one caused a given mistake.
Log-Mel spectrograms, the encoder–decoder transformer, the special task tokens that make transcription and translation the same model, why it sometimes hallucinates text over silence, and what really differs between tiny, base and small.
Why an MP4 is not a codec, the difference between soft and hard subtitles, SRT versus WebVTT in detail, the 23.976 fps sync problem, professional reading-speed limits, and what burning in subtitles actually costs.