How to Start Transcription Work for Clients
Typing from a blank page while audio plays is no longer the job. Software produces a first draft with timestamps and speaker guesses, and what a client pays a person for is everything that draft gets wrong.
That shift changed who this suits. Fast typing helps. Careful listening and stubbornness about detail matter more.
What the work is now that machines do the first pass
Run a recording through automatic transcription and you get something that reads fine until you check it. Names are wrong. Technical terms are wrong. Speaker labels swap in the middle of an exchange. Numbers are approximated. Punctuation lands in places that change what a sentence means, and where the audio was unclear the software produces a confident, plausible sentence that nobody said.
Your job is the correction pass: fixing names and jargon against a reference list, repairing speaker attribution, punctuating for meaning, marking genuinely inaudible sections with a timestamp instead of guessing, and formatting to whatever the client's template demands.
Some work still needs a human from the start. Recordings with heavy crosstalk, poor audio, strong regional accents, or a lot of specialist vocabulary defeat the software badly enough that correcting is slower than starting fresh.
The skill to build is proofreading against sound, not speed at a keyboard. You are checking a claim against evidence, sentence by sentence.
Clean verbatim versus full verbatim
Clean verbatim strips the noise: um, uh, false starts, repeated words, stammers, filler. What survives is what the person meant to say. Full verbatim keeps all of it, and adds markers for laughter, long pauses, overlapping speech and background events.
Clients care because they are doing different things with the file. A researcher analyzing how somebody speaks needs the hesitations, because the hesitation is the data. A company publishing a podcast transcript wants the version a human can read without wincing.
Ask which one before you touch the file, and ask for their style guide. Ask specifically about numbers, timestamps, how to label speakers, and what to do with a word you cannot make out.
If the client cannot answer, transcribe one minute both ways and send both with a short note asking which they want. That costs you a few minutes and saves you a full re-do, and it makes you look like someone who has done this before.
Audio you should refuse
Ask for a sample before quoting, always. Then listen to a minute from the middle of the recording rather than the opening — the opening is the clean introduction and it tells you nothing about the argument at minute forty.
Refuse or reprice: three or more people talking over each other, a recording made on a phone speaker across a table, a room with hard surfaces and audible echo, a session where the microphone was on one participant only, and material dense with terminology you cannot spell and cannot look up.
Accents deserve an honest answer rather than a brave one. If you cannot hold a speaker's rhythm at conversational speed, you will produce a file full of guesses, which is worse for the client than declining. Say you are not the right fit for this recording and offer to look at the next one.
Background music under speech is its own trap. It sounds tolerable in the sample and it exhausts you across an hour.
Headphones, foot pedals and what is genuinely optional
Headphones that seal properly are the one purchase that changes the work. Not expensive ones — sealed ones. Open-backed headphones and earbuds that leak leave you replaying the same three seconds.
A foot pedal is a real gain on long recordings, because your hands never leave the keyboard to rewind. On short clips it earns nothing. Treat it as something you buy after transcription has paid you, not before.
Genuinely optional: paid transcription software, a second monitor, a mechanical keyboard. What matters far more is knowing the playback shortcuts of whichever player you use, so that rewind, slow down and jump-back are muscle memory. Free players and free editors will do this job properly, and tools that cost nothing to start a business with covers what to install first.
A phone alone makes this work painful rather than impossible, and doing freelance work with only a phone is worth reading before you commit to a deadline on one.
Building speed without wrecking accuracy
Use passes instead of trying to be perfect in one sweep.
First pass: play the audio at normal speed while reading the machine draft, fixing as you go, and flagging anything uncertain with a marker you can search for later. Second pass: names, numbers, technical terms and speaker labels only, jumping between your flags. Third pass: read the text alone with the audio off, checking that sentences make sense as English. That last pass catches the confident invented sentence better than any amount of relistening.
Keep a glossary file per client — names, products, acronyms, place names, spelled the way that client spells them. Reuse it on their next job. Text expansion shortcuts for long recurring names save real minutes across a long file.
Do not raise playback speed to fix an earnings problem. Errors cost more time to find than the speed saved, and an accuracy complaint costs the client.
When the hours-per-audio-hour ratio stops being worth it
Time your first job honestly. Total hours worked, divided by the length of the audio. That ratio is the only number that tells you what this work actually pays you, and it is invisible when the job is quoted by the audio minute.
A clean single-speaker recording with a reference list gives one ratio. A four-person meeting in a café gives a much worse one, at the same quoted price, which is how people conclude transcription does not pay when what they took was the wrong file.
Categories where the ratio stays bad: multi-speaker meetings without a participant list, non-native speakers in a noisy environment, anything requiring you to research spellings, and legal or medical material where an error carries consequences and the checking is relentless.
Setting a rate per audio minute only works once you know your own ratio, so pricing work when you have no track record has to come after the measuring, not before it. Platforms differ enormously in what they list and what they take, and which freelance platforms are worth your time matters here more than in most services.
If your ratio does not improve across several jobs on similar audio, stop. This is service work with a hard ceiling, and the hours you are spending would build something better somewhere else.