Drop your audio file here
or browse
MP3, WAV, M4A, OGG — up to 100 MB
Voice memos, podcast files and recorded calls all work.
Paste a public Instagram Reel or video post URL
Turn an audio file into text — free, in your browser
Drop your audio file here
or browse
MP3, WAV, M4A, OGG — up to 100 MB
Voice memos, podcast files and recorded calls all work.
Paste a public Instagram Reel or video post URL
Audio to text — also described as transcribe audio, MP3 to text or voice to text — means converting the speech in a sound file into written words. The audio file is not changed; you get a text version of what was said, usually with a time marker at the start of each line so you can find that part of the recording again.
People run audio to text on voice memos and voice notes, interviews, podcast episodes, recorded meetings and lectures, dictation, and any clip where reading is easier than listening. Unlike dictation built into an app, this works on a file you already have.
Every line starts with the moment it was spoken, so long recordings stay easy to scan:
[00:00] Quick note to myself about the client call [00:06] they want the report by Friday morning [00:12] and they asked for the numbers broken out by region
Download the text as .txt to paste into notes or documents, or as .srt when you need timed captions.
MP3, WAV, M4A and OGG, and other formats your browser can decode. If it plays in the browser, the speech in it can be transcribed. Files up to 100 MB.
It is the same job described two ways: spoken voice in, written text out. A voice memo recorded on a phone is just an audio file, so it goes through the same box above.
No. The model runs in your browser and the file is not uploaded anywhere. There is no account to create and nothing is stored on a server.
Up to 100 MB per file. The audio is handled in 30-second passes, so an hour-long recording takes noticeably longer than a short memo and uses more memory in the browser.
Yes. Open this page in a mobile browser and pick the recording from your phone. The first run downloads the speech model, so allow a little extra time on mobile data.
English. The spoken language is identified first, and non-English audio returns a notice instead of a guess.
InstaScript is an independent personal project. It is not affiliated with, endorsed by, or sponsored by any platform or by OpenAI.