Audio Transcript Generator
Transcribe audio and video to text on your device, then copy it or download TXT or SRT subtitles.
The first transcript downloads a speech model once: about 40 MB for English, 80 MB for other languages. Your audio stays on your device.
How to transcribe audio to text
- Add an audio or video file: MP3, WAV, M4A, MP4, MOV, WebM and more.
- Pick the spoken language and click Transcribe.
- Keep the tab open while it runs. Long recordings take a while.
- Play the audio back and correct names, terms and punctuation.
- Copy the transcript, or download TXT or SRT subtitles.
Tips
- Transcription uses OpenAI's open-source Whisper model running in your browser. The first run downloads the model (about 40 MB for English, 80 MB for other languages) and caches it.
- Clear speech from one person at a time gives the best draft.
- A laptop or desktop transcribes much faster than a phone.
FAQ
Is my recording sent to a transcription service?
No. The speech model runs in your browser. The only download is the model itself, from the Hugging Face CDN.
How accurate is it?
Clear speech gives a useful draft. Expect to fix names, technical terms and punctuation.
Which languages are supported?
English uses a fast English-only model. Spanish, French, German, Italian, Portuguese, Dutch, Polish, Turkish, Russian, Ukrainian, Arabic, Hindi, Japanese, Korean and Chinese use a larger multilingual model.
Can I make subtitles?
Yes. Download SRT saves timed subtitles. Your corrections are kept if you leave one line per subtitle.
Why is the first transcript slow to start?
The model downloads once (40 to 80 MB). Later transcripts start right away.
Guides
- How Local Browser Editing Handles Your FilesHow client-side file editing works, how to verify local processing, and what page data is still shared.