Speech-to-text tool

Subtitle Extractor

Create timed subtitles from speech in your video. Everything happens in your browser.

  • No file upload
  • No sign-up
  • Free to use

Ready on your device

Choose a video you have the right to use. It is not uploaded.

Start here

Choose a video

Create timed subtitles from spoken English, Spanish, French, or Chinese.

Browser-local

Drop a video here

or

Use a video you own or have permission to process. Clear speech gives the best result.

Your video stays on your device.

Audio decoding and Whisper transcription happen in this browser. Your video is not uploaded or stored on a processing server.

A clear local workflow

How to create subtitles from a video

Choose a permitted video, identify the spoken language, and let Whisper work on your device.

  1. Choose a video you may use

    Select a video that you own or have permission to process. Choosing the file does not upload it. A preview and basic details appear after the browser reads the file. Clear speech with limited background noise generally produces the most useful transcript.

  2. Choose the spoken language

    Select English, Spanish, French, or Chinese. Whisper uses that choice to transcribe the original speech rather than translate it. The browser decodes the audio track and resamples it for the speech model without sending the video away.

  3. Review and save subtitles

    Review the timed text because names, numbers, accents, and noisy speech can be misheard. Save SRT for common editors and players, VTT for web video, or plain TXT when you only need the transcript.

What affects accuracy

Whisper Base can handle ordinary speech, but no automatic transcript is perfect. Quiet recordings, one speaker at a time, and clear microphones help. Music, overlapping voices, strong accents, technical names, low volume, and compressed audio can reduce accuracy. Always review subtitles before publishing them, especially when exact names, dates, prices, medical terms, or legal wording matter.

What local processing means

The Whisper model comes to your browser from NoUpload App's dedicated model host; your chosen video does not go to a transcription server. The first use downloads roughly 77 MB of ONNX model data plus the local runtime files. Your browser may cache these assets for faster reuse. Long videos need substantial memory and time, so begin with a short video while checking browser compatibility.

Straight answers

Subtitle extractor FAQ

Useful details before you transcribe a video.

Is my video uploaded?

No. The browser reads your video and decodes its audio locally. Whisper then creates the transcript in this tab. Selecting the file only gives this browser tab permission to read it; the video is not uploaded to us.

Why does the page download model files?

Speech recognition needs the quantized Whisper Base encoder and decoder. Those model files download to your browser before transcription. The transfer direction is toward your device; your selected video is not sent with those requests.

Which languages are supported?

The launch interface supports spoken English, Spanish, French, and Chinese. Choose the language actually spoken in the video. The tool transcribes that language and does not translate it into another language.

Which subtitle formats can I save?

You can save SubRip SRT, WebVTT VTT, or plain TXT. SRT works with many players and editors. VTT is designed for web video. TXT contains the spoken text without subtitle timing syntax.

May I transcribe any online video?

No. Use only files you own or have permission to process. This tool does not provide a downloader and is not intended to bypass copyright, platform rules, access controls, or another person's privacy.

Do I need an account or payment?

No account or sign-in is required, and the tool is free to use.