Subtitle Generator creates caption files from the speech in a video or audio recording. Drop in an MP4, WebM, MOV, MP3, or WAV file, and the Whisper speech model listens to the soundtrack on your own device and writes out every phrase with its start and end time. Download the result as SRT, the format YouTube Studio, Premiere Pro, DaVinci Resolve, Final Cut, CapCut, and VLC all import, or as WebVTT, the format HTML5 video players read through the track element. Nothing is uploaded: the video stays on your computer, which also means there is no file size quota, no watermark, and no account. Whisper places a cue roughly at every natural pause, so most cues run between two and eight seconds, which is close to the pacing subtitle editors aim for by hand. For foreign language footage, choose Translate to English and the subtitles come out in English while the timing still follows the original speech. A ten minute video typically produces 120 to 200 cues. If your editor cannot open the video format directly, Video Compressor re-encodes it to a standard MP4 first, and Audio Converter extracts a clean audio track.
Whisper is an encoder decoder transformer released by OpenAI in 2022 and trained on about 680,000 hours of multilingual audio collected from the web. That unusually broad training set is what makes it robust: it handles accents, background noise, and technical vocabulary far better than older speech recognizers trained on clean read speech. The model listens to audio in windows of exactly 30 seconds, converted into an 80 band log Mel spectrogram, and predicts text tokens together with timestamp tokens that mark where each phrase starts and ends. Longer recordings are cut into overlapping 30 second windows, each one transcribed separately, and the overlaps are then merged so words at the boundaries are not lost or duplicated. This tool uses windows that overlap by five seconds on each side, which is why the progress bar advances in steps. Running in the browser relies on ONNX Runtime compiled to WebAssembly, with the model weights quantized to 8 bit integers. Quantization shrinks the download by about four times compared with full precision weights, at a very small cost in accuracy. The two sizes offered here are the tiny model with 39 million parameters and the base model with 74 million. Larger Whisper models exist and are more accurate still, but at several hundred megabytes they are impractical to download into a web page. Audio is resampled to 16 kHz mono before recognition, because that is the input format Whisper was trained on; higher sample rates add nothing for speech. Subtitle output follows the two dominant standards: SRT, which uses a comma before milliseconds and numbered cues, and WebVTT, which starts with a WEBVTT header and uses a period.