Tools Audio Transcription
Audio Transcription

Transcribe your interview with each speaker separated and timestamped

Upload the recording, get the text split by speaker turn, rename each one, and export to Word or TXT.

Version
2026.08
Last updated
Available since

Free credits · No credit card · Instant access What changed recently

What this solves

Audio Transcription turns a recording into text divided by speaker turn, each block with a start and end time. Upload an MP3, WAV, M4A, AAC, OGG or WebM file up to 100 MB, say how many people speak or let the tool detect it, and the result comes back as Speaker 1 and Speaker 2, which you rename to the real names.

How it works

  1. Upload the recording

    Drag in an MP3, WAV, M4A, AAC, OGG or WebM file up to 100 MB. The audio language comes pre-selected as Portuguese; change it or switch to automatic detection.

  2. Say how many people speak

    Set the expected number of speakers, from one to six, or leave it automatic. Processing runs in the background, so you can close the page.

  3. Rename the speakers and export

    Read the transcript segment by segment with timestamps, replace Speaker 1 with the real name, then download a DOCX or TXT with everything.

Method

What happens to your file, step by step, and where the result stops being a decision made by a model.

How it works inside

The upload passes four checks: the extension must be one of the six allowed, the declared file type must match the extension, the first bytes of the file must look like real audio, and the size can be up to 100 MB. The job then runs in the background with a 30 minute ceiling. The default provider is AssemblyAI, asked to use the universal-3-pro model and fall back to universal-2, with speaker separation turned on; we check for the result every 5 seconds, for up to 600 seconds. Each stretch of speech becomes a segment with start and end in seconds; neighboring segments from the same speaker are merged and the labels renumbered from Speaker 1 up. If AssemblyAI fails with a temporary error, the audio is sent to OpenAI gpt-4o-transcribe-diarize instead.

Where AI is used, and where it is not

A speech model produces every word and every speaker boundary. No language model reads, rewrites, summarizes or corrects the transcript afterwards. Everything after the provider response is deterministic: language resolution (your choice, then the site locale, then auto-detect at 0.7 confidence), merging adjacent same-speaker turns, reassigning a ghost speaker (a single utterance under 5 seconds sitting between two turns of the same speaker), renumbering labels, and the credit arithmetic.

What it accepts

  • One audio file per request: .mp3, .wav, .m4a, .aac, .ogg or .webm. The extension, the declared MIME type and the first bytes of the file must all agree, or the upload is refused.
  • Hard cap of 100 MB per file. Empty files and files with fewer than 4 readable bytes are rejected.
  • Optional language: Portuguese (pre-selected in the form), English, Spanish, French, German, Italian, or automatic detection.
  • Optional Number of speakers field, from 1 to 6, passed on to the transcription provider. Optional title of up to 255 characters.

What you get back

  • Result page with the full text plus, per segment, the speaker label and the start and end time. Renaming a speaker rewrites the name on all their segments.
  • TXT export with a [HH:MM:SS - HH:MM:SS] Speaker header per segment, and DOCX with the same structure plus a metadata table (date, source, file, language, duration).
  • Stored metadata: provider, model actually used, requested and detected language, provider confidence, speaker count before and after consolidation, audio duration and per-phase timings.
  • Cost computed from the real duration: 2 credits per minute of audio, with the total rounded up to a whole credit and a minimum of 5 credits, charged before the record is marked complete. If the charge fails, the transcription is not delivered.

What this tool does not do

  • It does not name the speakers. Labels are Speaker 1, Speaker 2 and so on, assigned by order of first appearance; any real name is typed by you afterwards.
  • It does not summarize, translate, clean up or fact-check the audio. There is no minutes, topic extraction or highlight step of any kind.
  • It does not accept a URL and does not record inside the browser. The only way in is uploading a file in one of the six allowed formats; a video only passes if it is a WebM, and then only its audio track is transcribed.
  • It does not guarantee the language you picked. A mismatch between requested and detected language is logged and recorded in metadata, and the transcript is delivered anyway.
  • It does not export SRT or VTT subtitles and does not return word-level timestamps: the smallest unit is the utterance segment.
Plan
Pro
Cost
2 per minute (minimum 5)

What you get

  • Speaker separation

    Each turn is labelled Speaker 1, Speaker 2 and so on. Rename a speaker once and the new name shows in every segment and in the export.

  • Timestamps on every block

    Each block shows start and end time in hours, minutes and seconds, so you can go back to the recording and check a passage you doubt.

  • Runs in the background

    Long recordings keep processing after you leave the page. The transcript appears in your history when it is ready, with a status while it runs.

  • Word and TXT export

    Download a DOCX or TXT carrying the speaker names, the timestamps and the full text, ready to paste into your analysis.

Changelog

Every line below is a change that actually shipped, dated by the day it went out.

  1. Latest
    • The credit check now measures the real audio duration instead of guessing from file size.
    • A transcription is only delivered after the credits are charged, so nothing is handed over unpaid.
    • Failure messages now appear in the language you are reading the site in.
Show 4 earlier updates
    • Transcriptions saved with a project open are now collected inside that project.
    • Uploads are no longer refused for lack of balance when top-up credits cover the cost.
    • An invalid setting sent to the speech provider was fixed, so transcriptions stop failing there.
    • The tool page now requires login and a plan that includes transcription.
    • Interface texts and messages go through translation, so English readers see English.
    • Speaker separation and automatic language detection arrived with the new transcription provider.
    • The audio upload limit went from 50 MB to 100 MB.
    • The estimated cost of the transcription appears as soon as you pick the file.
    • A transcription stuck in processing is reported as failed instead of spinning forever.
    • First version of the tool: send an audio file, read the transcript, keep the history.
    • Long recordings are compressed and split automatically before being sent for transcription.
    • You can tell the tool how many speakers the recording has.
    • YouTube links existed in the first version and were removed the same day: only file upload works.

Questions and answers

Does it get every word right?

No automatic transcription does. Background noise, people talking over each other and strong accents cost accuracy, and the speaker split is an estimate. Treat the result as a draft you review against the audio.

What files and languages does it accept?

Audio files in MP3, WAV, M4A, AAC, OGG or WebM, up to 100 MB, uploaded from your computer. There is no link or video import. Language: Portuguese, English, Spanish, French, German, Italian or automatic detection.

Is my audio used to train AI models?

The file goes to the transcription provider on a paid API tier, whose published policy says API content is not used for training. We keep the transcript, not the audio. Our privacy policy lists every provider involved.

How much does a transcription cost?

Credits follow the audio duration: two per minute, with a minimum of five, charged after the transcript is ready. If the charge fails, nothing is delivered and nothing is taken. The tool is on the Pro plan.

How to cite this tool

Used it in your research? Here is the reference, already filled in with the version you are looking at and today's access date.

ABNT (NBR 6023)

LESSA, P. W. B. Audio Transcription. Versão 2026.08. [S. l.]: Xplore Dados, 2026. Disponível em: https://xploredados.com/en/tool/audio-transcription. Acesso em: 21 ago. 2026.

APA 7

Lessa, P. W. B. (2026). Audio Transcription (Version 2026.08) [Computer software]. Xplore Dados. https://xploredados.com/en/tool/audio-transcription

BibTeX

@software{xploredados_audio_transcription_2026,
  author  = {Lessa, Patrick Wendell Barbosa},
  title   = {Audio Transcription},
  organization = {Xplore Dados},
  version = {2026.08},
  year    = {2026},
  url     = {https://xploredados.com/en/tool/audio-transcription},
  urldate = {2026-08-21}
}

Start with the free credits

Create an account and test the tools before deciding on a plan.

Create free account