What the protocol usually needs to say
Most IRBs treat third-party transcription as a routine disclosure rather than an obstacle, but they expect it to be named. A protocol silent on transcription while the study plainly requires it invites a question you could have pre-empted.
The elements typically expected are: that a third-party service will be used, what safeguards apply to transmission and storage, whether identifiers appear in the recordings, how long the vendor retains material, and whether a confidentiality agreement is in place. None of this is onerous; it becomes a problem only when discovered mid-study.
Where recordings include identifiable health information and your institution is a covered entity, HIPAA obligations apply alongside the IRB requirement — they are separate regimes and satisfying one does not discharge the other. Clinical research sitting at that boundary carries both.
Verbatim convention is a methodological decision
This is the part most researchers discover too late, after transcripts arrive in the wrong convention and coding has already begun.
True verbatim retains filler words, repetitions, false starts, and non-verbal sounds. Conversation analysis and discourse analysis depend on these: a participant hesitating before an answer, or correcting themselves mid-sentence, is data. A cleaned transcript removes them for readability, which suits thematic analysis where the content of what was said matters more than its delivery.
The risk is silent. A transcript that smooths over hesitation and self-correction is not a slightly tidier transcript; it is evidence that has quietly changed. If your analysis treats hesitation as meaningful and your transcripts removed it, the finding is an artefact of the transcription convention. Specify which you need before files are sent.
- True verbatim — conversation analysis, discourse analysis, anything treating delivery as data.
- Cleaned verbatim — thematic analysis, content analysis, where meaning is carried by content.
- Speaker labels — required for focus groups if coding software is to distinguish participants.
- Timestamps — needed to cite a passage or return to the audio during analysis.
De-identification is harder in speech than in text
Protocols frequently specify that transcripts will be de-identified, and researchers often assume this is a find-and-replace operation applied after transcription. Spoken data resists that more than written data does.
Participants name people, employers, clinics, schools, and neighbourhoods spontaneously, in forms that do not match any list you prepared. They refer to a spouse by first name, mention the town they grew up in, or describe a workplace specifically enough to identify it without naming it. A transcript can contain no names at all and still identify a participant to anyone who knows the setting — which, in a study of a specific community or institution, includes exactly the readers most likely to see it.
Decide before transcription begins whether identifiers should be removed by the transcriptionist, flagged for you to remove, or left intact in a secure transcript that is de-identified later. All three are defensible; discovering that nobody chose is not. If the transcriptionist is de-identifying, they need instructions on what counts, because "remove identifiers" is not a specification.
Retention, and the copy you forgot about
An IRB protocol that specifies destroying recordings at study end has to account for every copy, and third-party transcription creates copies by design.
Ask three questions: how long does the vendor retain source audio after delivery, how long do they retain the finished transcript, and can you require deletion on request and get confirmation. A vendor retaining audio indefinitely for quality-assurance purposes is holding research data under your confidentiality obligations, past the date your protocol said it would be gone.
Grant and institutional requirements can pull the opposite way, requiring data to be preserved for a defined period after publication. Where those conflict with a protocol commitment to destroy recordings, the conflict is worth surfacing to your IRB before you sign with a vendor rather than after.
Focus groups and accented speech
Multi-participant recordings are the hardest material to transcribe accurately, and qualitative research produces them routinely. Speaker attribution in a six-person focus group with crosstalk is precisely where automated transcription degrades, and an attribution error in a focus group transcript propagates directly into coding.
Accented and non-native speech carries the same risk. Research populations are frequently more linguistically diverse than the corpora automated models were trained on, so error rates on study audio are often worse than published benchmarks suggest. Human review is where accented speech, code-switching, and unfamiliar proper nouns are corrected rather than guessed.
Sources
- 45 C.F.R. § 46 — Protection of Human Subjects (Common Rule)
- 45 C.F.R. § 164.308(b)(1) — business associate contracts, where PHI is involved
Verified 18 September 2026.
The regulatory information on this page is general background compiled from public primary sources, not legal or compliance advice. Requirements change and vary by jurisdiction and by court. Verify current rules with the relevant authority or your own counsel before relying on them.