How to Transcribe an Interview or Lecture (2026): Every Method, Priced in Hours and Dollars
If you have an hour of audio and a deadline, upload it to a service and you will have a rough transcript in a few minutes. If you have an hour of audio and an obligation to the person on it, that shortcut is the decision you should make slowly, because most of these services keep the file and some of them say plainly that they train on it. Everything below is organised around that fork: three routes that cost nothing, several that cost money, one that costs your afternoon, and an honest account of what each one gets wrong.
Transkriptor is the paid service that gets the whole job done for the least money: $99.99 for the year works out at $8.33 a month and covers forty hours of recording a month, with speaker labels and a written no-training promise on that plan.
If the recording must not leave your computer, use local Whisper through Buzz or MacWhisper. It is free, it handles twenty files as easily as one, and the audio never moves.
Our picks
| Role | Tool |
|---|---|
| Best free route if you already pay for Microsoft 365 | Transcribe in Word, 300 minutes a month |
| Best free route for a recording that must not leave your machine | local Whisper, through Buzz or MacWhisper |
| Best paid service for most people | Transkriptor |
| Best for meetings and interviews you host yourself | Otter |
| Best when the transcript has to hold up in court, on air, or in a citation | Rev human transcription |
| Best for one file with no subscription | Rev's AI transcription at $0.25 a minute |
| Best if the interview will become a podcast or a video | Descript |
| Best for multilingual teams | Notta |
| Best if you have decided to type it yourself | oTranscribe, free and open source |
One clarification before the methods, because the two words get mixed up constantly. Transcription turns an existing recording into text. Dictation turns your live speech into text as you type. If what you actually want is to talk instead of typing, you want a dictation app, and we have separate guides for Windows and Mac.
Before you upload anything
Answer one question and the rest of this page narrows to two or three options.
Did the person on the recording agree to have their voice sent to a company they have never heard of?
Agreeing to be recorded is not the same thing as agreeing to be uploaded. A source who says "sure, record it" is thinking about your notebook, not about a vendor's servers, a data labelling contractor, and a model training run. For a press conference or your own lecture notes this is a non-issue. For a whistleblower, a patient, a research participant who signed a consent form listing exactly who would handle their data, or anyone discussing something that could cost them their job, it decides the whole workflow.
If the answer is no, or you are not sure, skip to local Whisper. It runs on your own computer, costs nothing, and the audio never moves. Everything else on this page involves handing the file to somebody.
This is not a fringe worry. The Freedom of the Press Foundation, which trains journalists on exactly this class of decision, reviewed the popular transcription tools and reached a blunt conclusion in a piece updated on 2 June 2026: "each of these companies has the technical ability to access the audio you've uploaded", and "We recommend avoiding transcription altogether if your audio, in the wrong hands, could put people at risk." Their security trainers point the same readers at offline Whisper running on their own device.
A working researcher put the practical version of it better than we can. Writing on Hacker News on 17 July 2026, a developer explained why he had built his own transcription tool: "there are no real options for interview transcriptions with speaker detection that can be used in an university study evaluation context where you do not upload the data into a someone's else's cloud with questionable privacy policies."
At a glance
Costs are per hour of recording, because that is the unit you actually have. Everything checked 27 August 2026.
| Method | Best for | One hour of audio costs | Speaker labels | Where the audio goes |
|---|---|---|---|---|
| By hand | quotes you cannot afford to get wrong | 4 to 6 hours of your time | you write them | nowhere |
| Transcribe in Word | anyone with a Microsoft 365 subscription | free, within 300 min/month | yes | Microsoft |
| YouTube auto-captions | talks and lectures that are already public | free | no | |
| Local Whisper | confidential recordings, and bulk | free, plus machine time | only with extra tooling | nowhere |
| Whisper API | one batch, no software to install | $0.36 | no | OpenAI |
| Transkriptor | the whole job at the lowest price | $8.33/month covers 40 hours | yes | Transkriptor |
| Otter | live meetings you host | $8.33/user/month billed yearly | yes | Otter, which trains on it |
| Notta | multilingual meetings and long files | $97.99/year covers 30 hours a month | yes | Notta |
| Rev, AI | a single file, pay as you go | $15.00 | not stated | Rev |
| Rev, human | legal record, broadcast, published quotes | $119.40 | yes | Rev, plus a vetted transcriptionist |
| Descript | interviews that become audio or video | $16/month yearly for 10 media hours | yes, 8+ speakers | Descript, which trains on it unless you opt out |
Path 1. Typing it yourself
Best for: short clips, terrible audio, and material where you need to hear every hesitation.
Nobody in the software business wants you to know how long this takes, so here are three estimates from people who sell transcription and therefore have no reason to exaggerate the effort in your favour. SpeakWrite puts it at "about four times as long as the length of the audio". Otter's own guide says "3 to 6 hours per hour of audio". Happy Scribe's CEO, writing in September 2025, gives the working ratio as 4:1, "up to a 6:1 or even 8:1 ratio for complex transcriptions", with proofreading adding "another 25-50%" on top.
None of the three cites a study, so treat the numbers as what they are: a consistent trade estimate from people who do this daily. Call it a full working day for a one-hour interview, more if two people talk over each other.
Given that, why would anyone still do it? Three real reasons. Very short extracts are faster to type than to upload and clean. Audio recorded in a bar, on a phone, through a mask or in a strong accent can defeat every model on this page, and you will spend the day fixing the machine's guesses instead of writing your own. And in qualitative research, transcription is analysis: people who type their own interviews notice things they would never have caught skim-reading a clean transcript.
If you are doing it by hand, do not do it in a media player and a word processor. oTranscribe is a free, open source editor built for this one job: the audio player and the text sit in the same window, so you never alt-tab, and the keyboard controls pause and rewind without your hands leaving the keys. The project is MIT licensed, and its repository was still getting commits on 10 August 2026, which is more than can be said for most free tools of its generation. Your file stays on your computer.
Decide your style before you start, because changing halfway wastes the whole session.
- Verbatim keeps every "um", false start and repetition. Conversation analysis needs it. Nobody else does.
- Intelligent verbatim drops filler and stutters and keeps everything meaningful. This is the default for journalism and most research.
- Edited fixes grammar and half-sentences into readable prose. Good for a published Q&A, bad for anything you will later quote as speech.
Path 2. The free routes
All three are genuinely free, and none of the eight articles currently ranking for this topic mentions any of them. They also all have real limits, so read the failure conditions before you commit an afternoon.
Transcribe in Word
Best for: almost anyone who already pays for Microsoft 365 and has not noticed this feature exists.
Open Word in a browser, go to Home, click the arrow next to Dictate, choose Transcribe, and upload the file. Microsoft's own documentation sets the limits: "Users with a Microsoft 365 subscription can transcribe a maximum of 300 minutes of uploaded audio per month", and 30,000 minutes if you hold a Copilot licence. It takes .wav, .mp4, .m4a and .mp3, runs only in Edge and Chrome, needs a connection, and separates speakers automatically. Processing takes roughly as long as the recording. The result lands in a "Transcribed Files" folder on your OneDrive, and you can drop individual passages or the whole transcript straight into the document.
Five hours a month, free, with speaker separation, for a product tens of millions of people already pay for. That is the most underused thing in this category.
The catches, in order of how often they bite. It is cloud processing: Microsoft states that "Your audio files are sent to Microsoft and used only to provide you with this service", which is a clear commitment but still means the file leaves your machine. It only works in the browser version, not the desktop app. And a per-file size cap of 200 MB circulates widely, including in Microsoft's own support answers, but it is not on the current documentation page, so we are not going to state it as fact. Split long recordings if yours refuses to upload.
YouTube automatic captions
Best for: a talk, lecture or panel that is already on YouTube, or that you are happy to publish.
If the video exists on YouTube, you may already have a transcript. Click "Show transcript" in the description, and the text appears in a scrollable panel, timestamped, clicking any line to jump the video. If it is your own upload, YouTube Studio will give you the caption file itself: Subtitles, pick the video, Edit next to the language, Options, Download subtitles.
Google publishes the list of languages that get automatic captions, and it is long, running from Afrikaans to Zulu with around sixty-odd entries. Google also publishes, unusually plainly, when this fails. Captions do not appear for audio that is too complex to process, for unsupported languages, when "The video is too long", when there is "poor sound quality or YouTube doesn't recognize the speech", after "long period of silence at the beginning", and for "multiple speakers whose speech overlaps or multiple languages at the same time". That last one is a description of most interviews. And the accuracy warning is Google's own: "Automatic captions might misrepresent the spoken content due to mispronunciations, accents, dialects, or background noise."
There is a well-travelled trick of uploading a private or unlisted video purely to harvest the captions. It often works. We are not going to endorse it for two reasons: Google does not document whether privacy status affects caption generation, so the behaviour could change without notice, and uploading a source's interview to YouTube is uploading it to Google. For your own lecture recordings, fine. For anything sensitive, no.
Local Whisper
Best for: recordings that must not leave your computer, and for twenty files at once.
OpenAI released Whisper as an open model, and the ecosystem around it is now the strongest free option in this category. Nothing is uploaded, nothing is metered, and quality on clean audio is close enough to the paid services that most people stop noticing the difference.
The easiest way in is a GUI. Buzz is free and open source, runs on Mac, Windows and Linux, takes a file or a YouTube link, and exports plain text or subtitles. MacWhisper is the polished Mac option, freemium, and worth the money if you do this weekly (only download it from macwhisper.com, since the site itself warns that malware-carrying copycat sites exist). Underneath both sits whisper.cpp, which needs no dependencies, runs on CPU alone, and is heavily optimised on Apple Silicon, where Core ML can make the encoder "more than x3 faster compared with CPU-only execution".
Model choice is the one decision you have to make, and the trade is speed against accuracy. The large model is 1550M parameters and wants about 10 GB of VRAM. turbo is an optimised version of large-v3 that runs roughly eight times faster with, in OpenAI's words, "minimal accuracy loss". At the small end, tiny runs on anything and gets names wrong constantly. A practitioner's rule of thumb from Hacker News in June 2026: "medium is often the sweet spot for English accuracy vs speed, especially if following-up with a post-processing pass."
Set your expectations from people who run this daily rather than from the demo. As one of them wrote in August 2026: "It can work great as a first pass that someone can fix up, but not on its own." That is the honest summary of every route on this page, free or paid, and it is the reason none of them removes the editing step.
The genuine weakness is speaker labels. Plain Whisper gives you an undifferentiated wall of text, and it has done since 2023, when someone summed it up on Hacker News as "The only thing Whisper misses is speaker diarization". Three years on, the fix is still bolt-on tooling: WhisperX for word-level timestamps plus diarization, or whisper-diarization. Both work. Both add an install, a model download and a licence to read, and one developer moved off WhisperX in July 2026 specifically because of the licensing on the speaker model it depends on. For a two-person interview recorded on separate channels, you can sidestep the whole problem by transcribing each channel separately.
If you would rather not install anything, the same model is on OpenAI's API at $0.006 a minute, which is 36 cents for an hour of audio. That is a paid route, but it belongs here because it is the cheapest one on the page by two orders of magnitude, and because the cheaper gpt-4o-mini-transcribe sits at half that. You are back to uploading, of course.
The re-speaking workaround
Best for: nothing, honestly, but people keep asking about it.
The trick: open voice typing in Google Docs or Dictate in Word, play your recording out loud, and let the microphone pick it up. It half-works. Recognition is built for a live speaker at conversational distance, so a phone speaker across a desk drops words, punctuation collapses, and the session tends to stop when you switch windows. Google's documentation says nothing about recordings, playback or virtual audio devices, so nothing here is supported behaviour.
The version that does work is genuinely tedious: listen through headphones and repeat what you hear, clearly, into the microphone. Professional transcribers call it re-speaking and it is a real technique. It runs at roughly real time plus corrections, so figure ninety minutes for an hour. Given that Transcribe in Word does the same job unattended and free, this is now a fallback for people with no Microsoft 365 subscription and no ability to install software.
Path 3. The services
These do the job in minutes, format the result, label speakers and export something you can hand to someone else. The reason to pay is not accuracy; a local model gets close. The reason to pay is the formatting, the speaker labels and the export you would otherwise assemble by hand on every single file.
Transkriptor
Best for: getting the whole job done for the least money.
Price: a free plan with 90 minutes to start. Lite is $9.99 a month for 300 minutes. Pro is $19.99 a month, or $99.99 for the year, which works out at $8.33 a month and includes 2,400 minutes a month. Team is $30 per seat, or $240 a year.
Do that arithmetic and the position becomes obvious: Pro covers forty hours of recording a month for the price of a sandwich. If you are a graduate student with twenty interviews, or a journalist with a backlog of tape, no other service on this page is in the same range. It handles 100+ languages, identifies multiple speakers with editable labels, exports to TXT, SRT and Word (its interview page also lists PDF), and does the things you would otherwise do by hand: summaries, an AI chat over the transcript, translation.
Two things to be straight about. Its interview landing page claims "Achieve 99% Accuracy", while its lecture page says the same product "offers 99% accurate transcripts most of the time". The second version is the honest one, and it is still a vendor's number, not a measurement. And the line "Your data will not be used for training purposes" appears in the Pro feature list, not at the top of the site, which means the free and Lite plans carry no such promise. If you are uploading anything sensitive, that is an argument for paying rather than for trying it free.
Otter
Best for: meetings and interviews that you run and that you are recording live.
Price: the free Basic plan gives 300 minutes a month, capped at 30 minutes per conversation, and three file imports for the lifetime of the account. Pro is $16.99 a month, or $8.33 a month billed yearly, with 1,200 recording minutes, ten imports a month and a 90 minute cap per meeting. Business is $30, or $19.99 yearly, with unlimited meetings and four hours per meeting.
Read that free tier again, because every roundup lists Otter as a free way to transcribe an interview. It is not. Three uploads, ever. Otter is built to sit in a live call on Zoom, Meet or Teams and transcribe as people speak, and at that job it is very good: the text appears on screen during the conversation and corrects itself as context arrives. A deaf interviewer describing it on Hacker News called that live self-correction "a huge help", which is a use case no file-upload service can match. Feed it a recording from a handheld device, though, and you are using it against the grain.
Two further limits worth knowing before you commit. The pricing page names six languages: English, Spanish, French, German, Japanese and Chinese. And exports on Basic and Pro are mp3 and txt only. SRT, DOCX and PDF are Business features, so if you need subtitles or a formatted document, the real price is $19.99 a month, not $8.33.
Notta
Best for: multilingual work and long recordings.
Price: free gives 120 minutes a month with a three minute cap per recording, which makes it a demo rather than a plan. Pro is $97.99 a year, about $8.17 a month, for 1,800 minutes a month and up to five hours per recording. Business is $199.99 a year for unlimited transcription.
The five hour per-recording ceiling on Pro is the highest here and matters for anyone recording day-long sessions. Notta is SOC 2 Type II and ISO 27001 certified, does speaker identification and timestamped notes, and connects to Zoom, Meet, Teams and Webex. It comes out of Tokyo and is unusually strong in Japanese, which is the reason to pick it over Transkriptor if that is your working language. Its pricing page does not state a total language count or an export format list, and we could not find a statement anywhere on its site about whether recordings are used to train models. Absence of a claim is not a promise either way.
Rev
Best for: the transcript that has to be right, and for paying per file instead of per month.
Price: AI transcription is $0.25 a minute, which is $15.00 an hour of audio. Human transcription is $1.99 a minute, or $119.40 an hour. Subscriptions exist for volume: Essentials at $25.49 per seat a month billed yearly for 5,000 AI minutes, Pro at $47.99 for 10,000.
Human transcription is a different product from everything else on this page, and it deserves to be treated as one. Rev quotes "99%+ accurate, delivered in 12 hours or less" for it, against "95%+ accuracy in five minutes or less" for the AI version. That gap sounds small until you count it out. An hour of conversation runs to something like 9,000 words, so 95% leaves you around 450 wrong ones to find, and they cluster exactly where the meaning is: names, numbers and technical terms. If the transcript is going into a court filing, a broadcast, or a published quote a lawyer will read, $119 buys back an afternoon of checking and a category of risk.
The human side of the privacy story is the strongest of the paid options: transcriptionists "undergo rigorous vetting, including ID verification and NDAs", which is a real control and a rare one.
The machine side is where you have to read carefully, and it is the clearest example on this page of why marketing pages are not documents. Rev's security page says "we'll never train external LLMs on your data". Its terms of service, effective 15 May 2026, say something different: "Your Customer Content will be analyzed by our ASR models and other Rev artificial intelligence models and may be used for continuous training of those models". The exclusion carved out there is generative training, not training as such, and the contract itself offers no switch. An opt-out does exist, and it is easy to miss because it is not a setting: Rev's help centre, last updated 12 September 2025, says customers "are able to opt out of sharing their data for training purposes at any time, by emailing support@rev.com". An awkward route, but a written one. Both sentences are true at once, and the whole difference sits in the word external. Your interview will not end up inside someone else's foundation model. It may well end up inside Rev's.
Descript
Best for: an interview that will be cut into a podcast or a video.
Price: free gives one media hour a month. Hobbyist is $16 per person a month billed yearly, or $24 monthly, for 10 media hours. Creator is $24 yearly, or $35 monthly, for 30. Business is $50 yearly for 40.
Descript is not really a transcription service, which is why it is last. It is an editor where the transcript is the timeline: delete a sentence in the text and the audio goes with it, cut the filler words with one command, fix a misspoken phrase by typing over it. For an interview destined for publication as audio or video, that collapses two jobs into one. It transcribes 25 languages and detects 8 or more speakers.
Watch the unit. Descript bills media hours, which cover editing and export rather than transcription alone, so ten hours a month is not ten hours of interview if you also intend to edit them.
Its privacy policy does list "training our artificial intelligence models" among the purposes for which it processes content, but the control is better than that sentence suggests: training applies only to data you share through the "Share Data with Descript" setting. Which way that setting starts is less certain than it looks. The security page says it is disabled by default; the help centre describes the same control reading "Allowed" and tells you to click it to switch sharing off. The vendor's own documents do not settle the default, so check yours. What is clear is that a switch exists, which is more than an ordinary Otter account gets, and more convenient than Rev's, where the opt-out is an email. Credit where it is due.
What actually goes wrong
Every vendor on this page quotes a number between 95% and 99%. None of them publishes a test you can repeat. Here is what the independent research and the people doing this daily say instead, which is more useful, because errors are not distributed randomly. They land on exactly the words you were going to quote.
Accents. The most cited evaluation of Whisper across accents is Graham and Roll, published in JASA Express Letters in February 2024. Their finding, in their words: "Results reveal superior recognition in American compared to British and Australian English accents with similar performance in Canadian English", and "native English accents demonstrate higher accuracy than non-native accents". They also found error rates tracking with the speaker's first language and their proficiency in English. If you interview people who learned English as adults, your transcript will be measurably worse than the demo you tried, and the tool will not tell you so.
Names and jargon. This is the single biggest source of editing time, and it is the one thing no model can guess. A developer described it on Hacker News in July 2026: "Like OpenBao comes out as open bowel sometimes, lol. That necessitates significant cleanup." Substitute your source's surname, your field's acronyms and the name of the company you are writing about. Services that let you supply a custom vocabulary before processing will save you more time than a percentage point of headline accuracy.
Silence. This one surprises people. Speech recognition models can invent text where there is none. In "Careless Whisper: Speech-to-Text Hallucination Harms", presented at ACM FAccT in June 2024, researchers from Cornell, the University of Washington, NYU and the University of Virginia analysed more than 13,000 audio clips and found that roughly 1% of Whisper transcriptions contained entire hallucinated phrases. The trigger they identified was pauses: "Longer pauses and silences between words are more likely to trigger harmful hallucinations", which hit speakers with speech impairments hardest. The failure has its own folklore among practitioners, who recognise it by the stock phrases the model produces out of nothing. "The classic 'Thank you for watching!'", as one Hacker News commenter put it in February 2026; someone who had run thousands of hours of audio through Whisper reported the same thing arriving as "please like and subscribe". Both are artefacts of a model trained on a great deal of internet video, and both appear in silence.
There is a practical fix, and it is worth the ten minutes. Trim the silence before you transcribe, or run the audio through voice activity detection so the model is only ever fed speech. No pause, no invented sentence. Where trimming is not an option, the rule is simpler: if a line in your transcript reads oddly smoothly, or thanks somebody, check it against the audio before it goes anywhere.
Who said what. Speaker labelling is weaker than transcription everywhere, free and paid. Two voices of similar pitch, a crosstalk moment, someone joining halfway: labels drift, and they drift silently. Whatever tool you use, spot check speaker changes around every quote you intend to publish.
The practical consequence of all four: an AI transcript is a searchable draft, not a source. Every passage you plan to quote gets played back and checked against the audio. That final pass takes minutes rather than hours, and it is the step that separates a transcript you can work from and a transcript you can publish.
Who trains on your recording
We read these on 27 August 2026. They change, so the dates matter.
| Service | Trains on your recording? | What its own documents say | Document |
|---|---|---|---|
| Otter | Yes, no opt-out | "training our proprietary AI technology on de-identified audio recordings and on transcriptions (which may contain Personal Information)". Separately names "Data labeling service providers who provide annotation services and use the data we share to create training and evaluation data" | Privacy Policy, effective 16 June 2026 |
| Rev | Yes, opt out by email | "Your Customer Content will be analyzed by our ASR models and other Rev artificial intelligence models and may be used for continuous training of those models". The opt-out is not a setting: customers may opt out "by emailing support@rev.com" | Terms of Service, effective 15 May 2026; help centre, 12 September 2025 |
| Descript | Only on data you share | "training our artificial intelligence models" is listed among processing purposes, and it applies only to data shared through the "Share Data with Descript" setting. The security page calls that setting disabled by default; the help centre describes it reading "Allowed", so the starting state is not settled by Descript's own documents | Privacy Policy, 14 April 2025; security page and help centre |
| Transkriptor | No on Pro, as a written guarantee | "Your data will not be used for training purposes" is listed as a Pro plan feature, the only budget service here that puts a no-training promise in writing, and one more reason the paid tier is the right one for sensitive recordings. Retention: the window is yours to set, from one minute up; beyond that, data is kept "only as long as necessary", and "some Data may persist in backups" | Pricing page, Privacy Policy |
| Notta | Not stated | SOC 2 Type II and ISO 27001 certified. We found no statement about training on user recordings on its pricing, privacy or security pages | notta.ai |
| Local Whisper | No | nothing to state. The file does not move | n/a |
Three honest readings of this table.
"De-identified" is doing heavy work in Otter's sentence. Stripping account details from an audio file does not strip the voice, the names spoken inside it, or anything that was said.
The marketing page is not the document. Both Otter and Rev publish a reassuring line that is narrowly true and widely misread. Otter's privacy and security page says "No customer data will be used to train" its AI service providers' models, which is a promise about other companies, not about Otter. Rev's security page rules out external LLMs while its terms permit continuous training of Rev's own. In both cases the binding text is the one with an effective date on it.
And none of this is hidden. It is public, written in plain sentences, and not one article ranking for this topic quotes a word of it.
Consent to record and consent to upload are two different questions
Not legal advice. Recording law is genuinely complicated and jurisdiction-specific, so this is a signpost, not a ruling.
On recording, the reference journalists actually use is the Reporters Committee for Freedom of the Press. Its guide states that "Federal law requires the consent of at least one party before recording in-person, telephone or electronic conversations", under the federal wiretap statute at 18 U.S.C. §§ 2510 and 2511. On top of that, "About 11 states primarily have all-party consent requirements for recording. These states are California, Delaware, Florida, Illinois, Maryland, Massachusetts, Michigan (at least for recordings made by a third party who is not involved in the conversation), Montana, New Hampshire, Pennsylvania and Washington." A few states split the difference: Missouri and Oregon require everyone's consent for in-person conversations only, Connecticut and Nevada for phone calls only. For a call crossing state lines, the guide's advice is to "assume that the stricter state law will apply".
The hedging in that quote is deliberate. "About" and "primarily" are there because these laws carry exceptions, which is why the confident round numbers you will see in search snippets are worth less than the guide itself. Lawyers writing about AI notetakers in February 2026 add the operational point: read the vendor's terms, privacy policy and retention policy before the file goes up.
On uploading, there is mostly no law, and that is exactly why it needs a decision rather than a habit. Nothing stops you sending a lawfully recorded interview to a transcription vendor. Whether you should is a question about the promise you made to the person on the tape.
Three practical rules that cost nothing.
Say it out loud at the top of the recording. "I'm recording this, and I'll run it through an automated transcription service" takes four seconds, is on the tape forever, and turns an assumption into consent. If they hesitate, you have learned something important early.
Put it in the consent form if you are doing research, and check your own institution's list first. Ethics approval covers the data handling you described, and "a third-party cloud transcription service" is a data handling step. Recordings of participants count as identifiable data in their own right, which is how a university IRB will treat them regardless of what the transcript says. Institutions do not agree with each other here: Virginia Tech's libraries state that "Multiple transcription tools, including otter.ai and NVivo transcription, are not recommended by IT for Virginia Tech researchers", while other review boards accept a vendor's published privacy policy as sufficient. Which means the answer for you is a five minute question to your research office, not a judgement call from an article. Naming the vendor, the retention period and the deletion process is a paragraph of work at the proposal stage and an unfixable problem afterwards.
Keep one route that never leaves the building. Local Whisper takes an afternoon to set up once. Having it ready means the answer to "can I even upload this?" never has to be yes by default.
Getting a transcript you can use
Timestamps are the difference between a transcript and a document. Without them, verifying a quote means scrubbing through an hour of audio. Every paid service here does them; Rev syncs human-verified transcripts to the audio for every paragraph; Whisper produces them natively, and WhisperX takes them down to word level.
Speaker labels should be fixed before you edit anything else. Rename "Speaker 1" and "Speaker 2" to actual names in the first pass, while you still remember who sat where, and the whole file becomes searchable by person.
Pick your style once. Verbatim for conversation analysis, intelligent verbatim for almost everything else, edited only for a published Q&A. Switching midway creates a document nobody can quote from safely.
Export in the format your next tool wants. DOCX for editing and coding, SRT or VTT for subtitles, TXT for search and for feeding to something else, PDF only for sending to someone who will not edit it. Check this before you pay: Otter puts SRT, DOCX and PDF behind its Business plan.
Do the audio pass on quotes. Play back every passage you will publish. It is the fastest insurance on this page.
And the largest gain of all happens before any of this: record better. A phone lying on the table between two people, in a room without a fan or a coffee machine, beats every accuracy feature any vendor sells. If you interview regularly, a cheap clip-on microphone will improve your transcripts more than upgrading from a free plan to a paid one.
How we checked this
We did not run a hands-on accuracy benchmark, and we are not going to imply otherwise. Read the roundups above this one in the search results and you will find confident figures ("up to 99%", "tops out at 85%") published by companies that ranked themselves first, with no test file, no methodology and no raw data. Numbers you cannot check are worth less than no numbers.
What we did instead, on 27 August 2026.
- Read the current pricing page of every service here and took the prices and minute allowances off it directly, rather than copying figures from other articles.
- Read Microsoft's support documentation for Transcribe and Dictate, and Google's help pages for YouTube automatic captions and Docs voice typing, for the limits and the failure conditions quoted above.
- Read the privacy policies, security pages and terms of service of Otter, Descript, Rev, Transkriptor and Notta, and quoted them verbatim, which is how the gap between Rev's security page and Rev's contract came to light.
- Checked our conclusions against an independent authority rather than only against vendors: the Freedom of the Press Foundation's review of transcription tools, updated 2 June 2026.
- Checked the state of the open-source tooling through the GitHub API: whisper, whisper.cpp, WhisperX, Buzz, whisper-diarization and oTranscribe are all live and none is archived, with the most recent commits landing in August 2026.
- Took the accent and hallucination findings from the published research (JASA Express Letters, February 2024; ACM FAccT, June 2024) rather than from vendor marketing.
- Pulled dated practitioner comments from Hacker News through its public search API, so every quote here has a link and a date you can open.
Four things we could not verify, flagged rather than smoothed over. The full text of the JASA paper is paywalled from where we sit, so the accent findings come from its abstract. Transkriptor's pricing page does not say whether its free 90 minutes are a one-time allowance or monthly. Otter's enterprise documentation, which may describe an opt-out that ordinary accounts do not get, would not open for us. And Reddit was unreachable from our research environment, so community voices here come from Hacker News only, which skews technical.
One number we deliberately do not publish: how long it takes to clean up an AI transcript per hour of audio. Several articles quote one, none of them sources it, and the real figure swings from a few minutes on a clean solo lecture to most of an afternoon on a noisy three-way interview.
The competition
Tools we looked at and did not rank, with the reason.
Sonix, Trint, Happy Scribe, TranscribeMe, GoTranscript, Scribie and Temi are established and several are good. They occupy the same ground as the five services above at similar or higher prices, and adding them would have made the page longer without making the decision easier. Sonix and Happy Scribe in particular are worth a look if you need heavy multilingual output.
TurboScribe, ScreenApp, VEED and the browser-tool crowd rank heavily for "audio to text" and mostly resolve to a free tier with a short cap and a signup wall. If you have one short file and no account anywhere, they work. For an hour-long interview they are a worse deal than Word.
Meeting notetakers, meaning Fireflies, Fathom, Jamie and the rest, transcribe your calendar rather than your files. If your "interview" is a Zoom call you host, look at that category and at Otter. If it is a recording from a handheld device, they are not built for it.
MAXQDA, NVivo and ATLAS.ti include transcription inside qualitative analysis software. If you are already coding data in one of them, use what you have paid for. If you are not, do not buy analysis software to get a transcript.
Speechmatics, AssemblyAI and Deepgram are APIs for building this into your own product. Real quality, wrong shape for someone with a file and a deadline.
Dictation apps keep showing up in transcription roundups and should not. Wispr Flow, Voicy and the built-in Windows and Mac tools type your live speech. They do not process a recording you already have, which is the job on this page. If that is what you were after, start with dictating into ChatGPT or the Mac roundup.
What to do next
Take the first thirty minutes of your actual recording, the messy one with the air conditioning and the two people talking over each other, and run it through Transkriptor's free 90 minutes. Not a clean sample, not someone's demo file. Thirty minutes of your own worst audio tells you in five minutes whether any AI transcription is good enough for your accents, your names and your jargon, and that is the only question that decides whether you should pay for this at all. If the answer is no, you now know to budget for Rev or for an afternoon with oTranscribe. And if the recording should never have been uploaded in the first place, set up local Whisper instead and keep that route ready.
Frequently asked questions
How long does it take to transcribe a one-hour interview?
By hand, 4 to 6 hours for most people, and up to 8 if the audio is difficult or several people talk at once, plus another 25 to 50% for proofreading. With an AI service, processing takes minutes, and then you spend an unpredictable amount of time fixing names, terms and speaker labels. We are not going to put a number on that second part, because it depends entirely on your audio and nobody has published a figure worth repeating. With a human service like Rev, 12 hours of waiting and almost no work of your own.
What is the best free way to transcribe an interview?
If you pay for Microsoft 365, Transcribe in Word: 300 minutes a month, speaker separation included, nothing to install. If you do not, or if the recording is confidential, local Whisper through Buzz or MacWhisper. If the audio is already a public YouTube video, click "Show transcript" and you are done.
Can I transcribe audio in Microsoft Word?
Yes, in Word for the web only, using Home, Dictate, Transcribe. Microsoft 365 subscribers get 300 minutes of uploaded audio a month, in .wav, .mp4, .m4a or .mp3, in Edge or Chrome, with speakers separated automatically.
Is Otter really free?
Its free plan gives 300 minutes a month for live meetings you record inside Otter, but only three uploaded file imports for the entire life of the account. If your interview is a file from a recorder, the free plan will not do the job more than three times.
Do I need permission to upload a recording to a transcription service?
Legally, in most places, no separate permission is required beyond the consent you needed to record. Ethically, and in any research governed by an ethics committee, yes: the person consented to being recorded by you, not to their voice being processed by a third party that may keep it and train on it. Say it on the tape, or write it into the consent form. The Freedom of the Press Foundation goes further for high-risk material and recommends "avoiding transcription altogether if your audio, in the wrong hands, could put people at risk."
Is it legal to record an interview?
It depends where everyone is. Federal law in the US requires the consent of at least one party. About eleven states require consent from everyone, including California, Florida, Illinois, Maryland, Massachusetts, Pennsylvania and Washington. On calls between states, assume the stricter law applies. This is a signpost, not legal advice, and the Reporters Committee for Freedom of the Press keeps the state-by-state detail.
How accurate is AI transcription really?
On clean audio in American English, good enough that most of your editing will be punctuation and names. Accuracy drops measurably on non-native and non-American accents, drops further on overlapping speech, and fails in a specific way on proper nouns and technical vocabulary. Models can also insert text that was never spoken, particularly during silences: peer-reviewed work found this in about 1% of Whisper transcriptions. Treat any transcript as a draft and check quotes against the audio.
Which tools label speakers?
Transkriptor, Otter, Notta, Descript (8 or more) and Word's Transcribe all do it automatically. Rev's human transcripts do it reliably. Plain local Whisper does not, and you need WhisperX or a similar add-on to get it.
How do I transcribe a lecture recorded on my phone?
Transfer the file to a computer and upload it to Transcribe in Word if you have Microsoft 365, or run it through Buzz locally. If the lecture is longer than five hours, check the per-recording cap before you pay for anything: Notta's Pro plan handles five hours per file, several rivals stop sooner.
What is the difference between transcription and dictation software?
Transcription turns a recording you already have into text. Dictation turns your live speech into text in whatever application you are using. They are different products, and tools that are excellent at one are usually useless at the other.
Comments
No comments yet. Ask a question or share what worked for you.