Speech to Text Audio Transcriber
About Speech to Text Audio Transcriber
Speech to Text Audio Transcriber turns recordings into text using Whisper, the open source speech recognition model from OpenAI, running directly inside your browser. Drop in an interview, a lecture, a voice memo, a podcast episode, or a video, and the tool decodes the audio, runs it through the model on your own device, and returns a transcript with timestamps for every segment. Your recording is never uploaded, which matters for meetings, medical or legal dictation, and anything else you would not hand to a free web service. Two model sizes are available: Fast is about 40MB and handles clear speech well, while Accurate is about 80MB and copes better with accents, background noise, and fast talkers. Either one is downloaded once and then cached. Whisper understands dozens of languages and can detect the language automatically, and it can also translate speech from any supported language straight into English text. Results can be copied or downloaded as plain text, as SRT subtitles for video editors and YouTube, or as WebVTT for web players. A second mode, Live dictation, types what you say into the page in real time using your browser speech recognition. For the opposite direction, Text to Speech Converter reads text aloud, and Audio Converter converts unusual recording formats to MP3 or WAV first.
Whisper is an encoder decoder transformer released by OpenAI in 2022 and trained on about 680,000 hours of multilingual audio collected from the web. That unusually broad training set is what makes it robust: it handles accents, background noise, and technical vocabulary far better than older speech recognizers trained on clean read speech. The model listens to audio in windows of exactly 30 seconds, converted into an 80 band log Mel spectrogram, and predicts text tokens together with timestamp tokens that mark where each phrase starts and ends. Longer recordings are cut into overlapping 30 second windows, each one transcribed separately, and the overlaps are then merged so words at the boundaries are not lost or duplicated. This tool uses windows that overlap by five seconds on each side, which is why the progress bar advances in steps. Running in the browser relies on ONNX Runtime compiled to WebAssembly, with the model weights quantized to 8 bit integers. Quantization shrinks the download by about four times compared with full precision weights, at a very small cost in accuracy. The two sizes offered here are the tiny model with 39 million parameters and the base model with 74 million. Larger Whisper models exist and are more accurate still, but at several hundred megabytes they are impractical to download into a web page. Audio is resampled to 16 kHz mono before recognition, because that is the input format Whisper was trained on; higher sample rates add nothing for speech. Subtitle output follows the two dominant standards: SRT, which uses a comma before milliseconds and numbered cues, and WebVTT, which starts with a WEBVTT header and uses a period.
How to use Speech to Text Audio Transcriber
- Drop in an audio or video file, or switch to Live dictation to use your microphone.
- Pick the Fast or Accurate model, the spoken language, and whether to translate to English.
- Click Transcribe, then copy the text or download it as TXT, SRT, or VTT.
Frequently Asked Questions
- Is my audio uploaded to a server?
- No, not in the Transcribe a file mode. The Whisper model is downloaded to your browser and the transcription runs on your own CPU, so the recording never leaves your device and nothing is stored anywhere. You can confirm this in your browser network tab: after the one time model download, no audio data is sent. Live dictation is different, because Chrome and Edge send microphone audio to their online speech service.
- How accurate is the transcription?
- On clear speech with one speaker, Whisper gets the large majority of words right, and the Accurate model does noticeably better than Fast with accents, background noise, and overlapping speech. Names, jargon, and very quiet passages are the usual weak spots. Always give an important transcript a quick read before relying on it.
- How long does it take to transcribe a file?
- It depends heavily on your hardware. On a recent laptop the Fast model usually runs faster than real time, while the Accurate model is roughly half as fast. Phones and older computers are slower. Text appears on screen as it is recognized, and keeping the tab in the foreground avoids browser throttling.
- Which file formats can I transcribe?
- Anything your browser can decode: MP3, WAV, M4A, AAC, OGG, Opus, FLAC, and WebM audio, plus the soundtrack of MP4 and WebM videos. Safari and Chrome cover slightly different sets of formats. If a file cannot be read, convert it to MP3 or WAV with the Audio Converter and try again.
- Can I create subtitles for a video?
- Yes. Drop the video file in, transcribe it, and download the result as SRT or VTT. SRT is accepted by YouTube, Premiere Pro, DaVinci Resolve, CapCut, and most video players, while VTT is the format used by HTML5 video on the web. Each subtitle line keeps the start and end time Whisper detected.
- What languages are supported?
- Whisper was trained on 99 languages, and the language menu lists 34 of the best supported ones, including English, Spanish, French, German, Italian, Portuguese, Russian, Polish, Serbian, Croatian, Turkish, Arabic, Hindi, Chinese, and Japanese. Auto-detect works well for most recordings; choosing the language manually helps with short clips and heavy accents. The Translate to English option turns speech in any of these languages into English text.
- Can I transcribe a Zoom, Teams, or Google Meet recording?
- Yes. Meeting recordings are ordinary MP4 or M4A files, so drop the file in and transcribe it like any other recording. Because nothing is uploaded, this is a safe option for internal meetings. Whisper does not label who is speaking, so the transcript is one continuous text with timestamps rather than a list of named speakers. The first run downloads the 40 to 80MB model once; after that it is cached.
Related Tools
Also Available As