Start with a short recording
Open Audio to Text in desktop Chrome and choose a clear recording. Set the language actually spoken. Base is the default; tiny uses a smaller model but can make more mistakes. Small is an optional upgrade with a larger download, higher memory use and usually longer processing. The current limits are 200 MB and 20 minutes, not a guarantee that every device can process that much.
Understand the first download
The model is downloaded before speech recognition. Base weights are about 206 MB for GPU processing or 77 MB for CPU processing, plus roughly 21 MB of runtime and configuration files. Browser cache can avoid another model download, but clearing storage or switching model versions can require it again. Keep the tab open during recognition.
Know what stays local
Your chosen recording and generated text stay in the browser. The website and public model files still require network requests. This is not a promise of complete offline availability. Export TXT or subtitles before closing the page; there is no cloud project recovery.
Example
Choose a 10-second recording → select its language → start → check the transcript → download TXT