Skip to content

Free while in beta - no account needed

Turn recordings into accurate text in minutes

SautiScribe transcribes interviews, meetings, lectures, podcasts and videos into clean, timestamped transcripts. It detects the language automatically across more than 90 languages, Swahili included, and gives you subtitles you can use straight away.

Max file size
100 MB
Max length
15 min
Languages
90+
Auto-deleted after
24 h

How it works

Three steps, no software to install.

  1. 1

    Upload your recording

    Drag in an audio or video file up to 100 MB. It uploads directly to encrypted storage, so nothing is held on our servers along the way.

  2. 2

    We transcribe it

    A GPU picks up the job and runs WhisperX. You watch live progress as it detects the language, transcribes the speech and aligns the timestamps.

  3. 3

    Read and download

    Read the transcript in your browser with timecodes, then download it as plain text, or as SRT or VTT subtitles ready to drop into a video editor.

Built for real recordings

Automatic language detection

Over 90 languages, identified from the audio itself. Override it manually when a recording is short, noisy or code-switches between languages.

Swahili and East African languages

Whisper transcribes Swahili, Somali, Amharic and Hausa. Timing accuracy varies by language, and SautiScribe tells you when timestamps are approximate rather than pretending otherwise.

Three output formats

Plain text for reading and quoting, SRT for video editors, and WebVTT for the web. All generated from the same aligned transcript.

Audio or video, no conversion

Upload MP3, M4A, WAV, MP4 or MOV. If there is an audio track in the file, the audio is extracted for you.

Deleted automatically

Your audio and transcript are removed within 24 hours. Nothing is used for training, and nothing is passed to third parties.

Nothing to install

It runs in the browser. No desktop app, no plugins, no account, and no waiting for a subscription to activate.

Frequently asked questions

What is SautiScribe?

SautiScribe is a web app that turns audio and video recordings into accurate, timestamped text. Upload a file, wait a couple of minutes, then read the transcript in your browser or download it as TXT, SRT or VTT subtitles. No account is required.

Which languages does it support?

Over 90, detected automatically. That includes English, Swahili, Somali, Amharic, Arabic, French, Spanish, Portuguese and Hindi. You can also set the language manually if a recording is short, noisy, or mixes languages, which is where automatic detection tends to struggle.

How accurate is the transcription?

SautiScribe runs WhisperX, which combines OpenAI's Whisper large-v3 model with forced alignment for precise segment timings. Accuracy depends on the recording: clear speech on a decent microphone transcribes very well, while heavy background noise, distant speakers or crosstalk will produce mistakes you should proofread.

What file formats can I upload?

MP3, M4A, WAV, MP4 and MOV, plus WebM, OGG and AAC. If your file has an audio track, SautiScribe extracts it automatically - you do not need to convert video to audio first.

What happens to my files?

Your audio is uploaded straight to encrypted object storage, transcribed, and then deleted automatically within 24 hours along with the transcript. Nothing is used to train models and nothing is shared with third parties.

Is SautiScribe free?

Yes, within limits. Transcription runs on GPUs that cost real money, so anonymous use is capped per file length and per day. Larger allowances and saved transcript history are planned.

Ready to transcribe something?

Upload a recording and read the transcript a couple of minutes later. No sign-up, no card.

Transcribe a file