← Voxoira home

Audio to text converter — free and private

Upload an audio or video file and get an editable transcript with timestamps. Transcription runs on your device using Whisper, so your file is never uploaded.

Choose a file to begin.

How to convert audio to text

Why use a browser-based transcriber

Many online converters upload your recording to a server, limit free minutes, or require an account. This tool processes the file locally, which suits interviews, lectures, voice notes and private meetings. Because it runs on your own computer, long files take longer on slower machines, and closing the tab stops the job. Use the Base model for better accuracy and keep the tab open until it finishes.

Making subtitles

SRT and VTT exports include a start and end time for each segment, so you can load them into video editors and players. Timing follows the speech segments Whisper detects, so check long files and adjust line breaks where needed.

FAQ

Does it detect different speakers automatically?

No. Speaker labels are manual: rename up to four speakers and mark where each one starts talking. Labels are included in TXT, SRT and VTT exports.

Is my audio uploaded to a server?

No. The file is decoded and transcribed in your browser. Only the speech model itself is downloaded once from a public CDN and then cached.

Which file types work?

Common audio and video files such as MP3, WAV, M4A, MP4 and WebM, depending on what your browser can decode.

How accurate is it?

Whisper is strong on clear speech. The Tiny model is fastest but least accurate; Base is slower and better. Noisy or overlapping speech will need editing.

How long does it take?

It runs on your device, so speed depends on your computer. Expect a few minutes for a 10-minute file, plus a one-time model download.

Can I get subtitles?

Yes. Export SRT or VTT files with timestamps, or plain TXT.

Specific formats: MP3 · M4A · WAV · Voice message

Need live dictation instead? Try voice to text or the voice translator.