Audio to Text

Turn interviews, lectures, and voice notes into editable text with local Whisper transcription.

Local fileText & subtitles

Drop your audio or video here

MP3 · WAV · M4A · WEBM · OGG · OPUS · AAC · FLAC · Up to 200 MB / 20 minutes

Processed in your browser. No file upload.No signup · Free exports

What can Audio to Text do?

Choose a recording, set its spoken language, and review the words before exporting. Whisper transcribes speech in its original language. Use the result to search a lecture, check an interview quote, or prepare a transcript. It does not summarize meetings or identify speakers.

How to use Audio to Text

Prepare your file, check the result, and download the format you need.

  1. 01

    Prepare your file

    Choose a clear MP3, WAV, M4A, OGG, OPUS, AAC, or FLAC recording that your browser can decode.

  2. 02

    Process and review

    Select the spoken language and start local recognition. The model downloads on first use.

  3. 03

    Download your result

    Check names and numbers, then download TXT text or SRT/VTT subtitles.

Formats & example

MP3 · WAV · M4A · WEBM · OGG · OPUS · AAC · FLAC → TXT / SRT / VTT

interview.mp3 → review text → interview.txt

Protective limits: 200 MB and 20 minutes. Tested in desktop Chrome. Device performance and codecs affect processing; try a short clip first.

Your files stay in this browser

Whisper recognizes speech on your device. The first use downloads the program and model. Your files and transcript are not uploaded. Results are not saved automatically. Download before leaving.

Audio to Text: common questions

Useful answers before you begin.

Is my recording uploaded?

No. Recognition runs in your browser. Network requests download the website and public model files. Export your result before leaving; there is no cloud project storage.

How long can my recording be?

The protective limits are 200 MB and 20 minutes. Actual capacity depends on your device. Try a short clip first; the limits are not a guarantee for every phone.

What about noise or accents?

Noise, overlapping voices, accents, and specialist terms can cause errors. Use a clear source recording and listen back to important passages.