Voice or Speech? Fifty Years of Transcription and the Rise of AI-Based Transcription Systems
Plain Language Summary
Turning a recorded conversation into a written transcript looks like clerical work, but every transcript is the product of choices — whether to keep the "ums," how to mark a pause, how to lay speakers out on the page — and those choices quietly shape what a study is able to find. Automatic speech recognition, the technology behind the transcripts that meeting platforms now generate on their own, makes many of those choices on the researcher's behalf: quickly, cheaply, and largely out of sight. We trace half a century of thinking about transcription, review the evidence that these systems hear some voices considerably better than others, and test whether the notation a transcript uses changes how well an AI language model can interpret it. Our aim is to return transcription to view as something researchers decide rather than something that simply arrives.
Contribution
We synthesize fifty years of transcription scholarship with contemporary evidence of ASR bias to argue that AI-based transcription does not resolve Ochs' "transcription as theory" but defers and obscures it, and we give that argument empirical purchase by showing that Jefferson notation measurably improves LLM interpretation of a transcript within what we term the "transcript-recoverable boundary."
Implications
Because a researcher cannot code a transcript they do not have, and cannot have one without something or someone transcribing, transcription remains a vital touchpoint between qualitative researchers and the people, objects, and contexts of their investigation even where it appears minimized and distanced; treating it as a solved problem, or relegating it to a "fetish," would be a grave mistake for the field. Upholding research quality accordingly requires that researchers describe their transcription processes in publications — how recordings were obtained, which ASR systems were used and on what justification, which notation system and with what modifications, how transcripts were verified and corrected, and what the act of transcribing surfaced. Where LLMs enter the workflow, that disclosure extends to how notation was matched to the model, since the context a transcript carries demonstrably shapes what these systems can interpret.
