Transcribe audio and video to text
Podcasts, interviews, voice notes, lectures and meetings — upload the recording and get the text back, with timestamps when you need them.
Drop a file hereor click to choose one — it opens in the studio with this tool ready. Small files need no account.Choose a fileHow it works
- 1Upload an audio or video file — MP3, WAV, M4A, OGG, FLAC, MP4, MOV and more.
- 2Choose the output: plain text, subtitles (SRT, VTT) or JSON with timings.
- 3Download the transcript.
What you get
Audio and video
Voice notes, podcasts and phone recordings go in as they are — no need to convert them to video first.
Word-level timestamps
The JSON output carries the start and end of every word, ready for search, highlighting or your own editor.
Your names and terms
Over the API, pass a list of product names, people and jargon so recognition spells them the way you do.
Several formats in one go
Ask for text, SRT and VTT in one request and get them together in one archive.
Same thing over the API
One header, one multipart request with the file, then poll the job or take a webhook. The same operation you just tried in the browser.
Questions
How accurate is it?
It uses OpenAI Whisper. Clear speech in a major language transcribes very well; setting the spoken language explicitly helps with accents and short clips.
Can it translate while transcribing?
Yes, into English: choose “Translate → English” as the result language.
What happens if there is no speech?
You get an empty transcript, not an error — useful when you process many files automatically.
Can I send files from my own app?
Yes. The transcription API takes the file in one request and returns a job you poll, or calls your webhook when it is done.