Automatic language detection
Over 90 languages, identified from the audio itself. Override it manually when a recording is short, noisy or code-switches between languages.
Free while in beta - no account needed
SautiScribe transcribes interviews, meetings, lectures, podcasts and videos into clean, timestamped transcripts. It detects the language automatically across more than 90 languages, Swahili included, and gives you subtitles you can use straight away.
Three steps, no software to install.
Drag in an audio or video file up to 100 MB. It uploads directly to encrypted storage, so nothing is held on our servers along the way.
A GPU picks up the job and runs WhisperX. You watch live progress as it detects the language, transcribes the speech and aligns the timestamps.
Read the transcript in your browser with timecodes, then download it as plain text, or as SRT or VTT subtitles ready to drop into a video editor.
Over 90 languages, identified from the audio itself. Override it manually when a recording is short, noisy or code-switches between languages.
Whisper transcribes Swahili, Somali, Amharic and Hausa. Timing accuracy varies by language, and SautiScribe tells you when timestamps are approximate rather than pretending otherwise.
Plain text for reading and quoting, SRT for video editors, and WebVTT for the web. All generated from the same aligned transcript.
Upload MP3, M4A, WAV, MP4 or MOV. If there is an audio track in the file, the audio is extracted for you.
Your audio and transcript are removed within 24 hours. Nothing is used for training, and nothing is passed to third parties.
It runs in the browser. No desktop app, no plugins, no account, and no waiting for a subscription to activate.
SautiScribe is a web app that turns audio and video recordings into accurate, timestamped text. Upload a file, wait a couple of minutes, then read the transcript in your browser or download it as TXT, SRT or VTT subtitles. No account is required.
Over 90, detected automatically. That includes English, Swahili, Somali, Amharic, Arabic, French, Spanish, Portuguese and Hindi. You can also set the language manually if a recording is short, noisy, or mixes languages, which is where automatic detection tends to struggle.
SautiScribe runs WhisperX, which combines OpenAI's Whisper large-v3 model with forced alignment for precise segment timings. Accuracy depends on the recording: clear speech on a decent microphone transcribes very well, while heavy background noise, distant speakers or crosstalk will produce mistakes you should proofread.
MP3, M4A, WAV, MP4 and MOV, plus WebM, OGG and AAC. If your file has an audio track, SautiScribe extracts it automatically - you do not need to convert video to audio first.
Your audio is uploaded straight to encrypted object storage, transcribed, and then deleted automatically within 24 hours along with the transcript. Nothing is used to train models and nothing is shared with third parties.
Yes, within limits. Transcription runs on GPUs that cost real money, so anonymous use is capped per file length and per day. Larger allowances and saved transcript history are planned.
Upload a recording and read the transcript a couple of minutes later. No sign-up, no card.
Transcribe a file