No. 02

Audio/Video Transcriber

Turn speech in an audio or video file into text, with SRT/VTT subtitle export — transcribed on your device with Whisper.

This browser doesn't support the Web Audio API needed to decode audio locally.

Uploaded 0 bytes Loads a Whisper speech model (100–300MB depending on quality) once, then works offline. (network monitor unsupported in this browser)

How your data is handled: everything above runs inside this browser tab; your file is never uploaded. Loads a Whisper speech model (100–300MB depending on quality) once, then works offline.

Sending a recording to a transcription service means trusting a third party with everything said in it. This tool runs OpenAI’s Whisper model directly in your browser, so the audio and the resulting transcript never leave your device.

How to use it

  1. Drop an audio or video file onto the waveform.
  2. Pick a model size: faster/smaller or slower/more accurate.
  3. Click Transcribe and watch the transcript scroll in sync with the waveform as it processes.
  4. Export as plain text, SRT, or VTT subtitles.

Tips & edge cases

  • The smallest model is noticeably faster but makes more mistakes on accents, background noise, and technical vocabulary — use a larger model when accuracy matters more than speed.
  • Background music or overlapping speakers reduce accuracy more than audio quality does; a clean single-speaker recording transcribes best regardless of model size.
  • WebGPU (when available) is several times faster than the WASM fallback — if transcription feels slow, check whether your browser and GPU support it.

FAQ

Is my audio or video uploaded to transcribe it?
No. Transcription runs Whisper entirely in your browser via transformers.js — WebGPU when available, WASM otherwise. Nothing is sent over the network during processing.
Why does it need to download a model first?
The speech-recognition model itself is 100–300MB depending on the quality tier you pick. It's downloaded once and cached, so later visits and later files skip that step.
Which languages are supported?
Whisper supports dozens of languages; recognition quality varies by language and is generally strongest for widely-spoken ones with more training data.