Audio Transcript Generator

Transcribe audio and video to text on your device, then copy it or download TXT or SRT subtitles.

Add a recording to transcribe: drop an audio or video file here or click to browse
MP3, WAV, M4A, AAC, OGG, FLAC, and WebM audio when supported by your browser. Files stay on your device.

The first transcript downloads a speech model once: about 40 MB for English, 80 MB for other languages. Your audio stays on your device.

Add an audio or video file to start.
Advertisement

How to transcribe audio to text

  1. Add an audio or video file: MP3, WAV, M4A, MP4, MOV, WebM and more.
  2. Pick the spoken language and click Transcribe.
  3. Keep the tab open while it runs. Long recordings take a while.
  4. Play the audio back and correct names, terms and punctuation.
  5. Copy the transcript, or download TXT or SRT subtitles.

Tips

FAQ

Is my recording sent to a transcription service?

No. The speech model runs in your browser. The only download is the model itself, from the Hugging Face CDN.

How accurate is it?

Clear speech gives a useful draft. Expect to fix names, technical terms and punctuation.

Which languages are supported?

English uses a fast English-only model. Spanish, French, German, Italian, Portuguese, Dutch, Polish, Turkish, Russian, Ukrainian, Arabic, Hindi, Japanese, Korean and Chinese use a larger multilingual model.

Can I make subtitles?

Yes. Download SRT saves timed subtitles. Your corrections are kept if you leave one line per subtitle.

Why is the first transcript slow to start?

The model downloads once (40 to 80 MB). Later transcripts start right away.

Guides