Open a captioning app today and something has changed in how it works. The text does more than trail the speaker word by word now. It guesses ahead, fills gaps from context, and corrects itself a beat after it gets something wrong. The newest AI captioning is genuinely better than it was even two years ago. It is also still wrong in the places where being wrong matters most. Both of those things are true, and knowing the difference is the whole game.
What “predictive” captioning actually does
Older speech-to-text worked roughly one word at a time. Newer systems use large language models that hold the whole sentence in mind, so they use context to choose between words that sound alike and to repair a phrase once the rest of the sentence makes the meaning clear. You may have seen the text on screen briefly change after it first appears. That is the model revising its first guess against everything it has heard since.
This is why captioning feels smoother now. The system does more than hear sounds. It models what a sentence is likely to be. When the speech is clear and the topic is ordinary, the result is fast, readable, and often very accurate. Be careful how far you push that, though. OpenAI’s own paper on Whisper, the model behind a lot of current captioning, claims something narrower than a leap in raw accuracy. Trained on 680,000 hours of speech, it says the models are often competitive with earlier fully supervised systems without any fine-tuning for the task at hand, and that they approach human accuracy and robustness. The gain is in holding up across accents, noise, and subject matter it was not tuned for. That is a real improvement, and it is not the same as being right more often in the hard rooms.
No spam. No inspiration porn.
Our best writing for adults with disabilities, weekly and free.
Get the newsletterWhere it genuinely helps
For everyday access, the gains are concrete. Google Live Transcribe is free, and Google Research says it partnered with Gallaudet University on the user research behind it. It handles clear speech well even with some background noise, and it shows you the volume of the speaker’s voice against the background so you can tell when the microphone is losing the room. Otter.ai will join a Zoom, Teams, or Google Meet call and caption it live, then hand you a searchable transcript afterward. For a university lecture, a staff meeting, a webinar, or a one-on-one conversation in a reasonably quiet room, these tools can be the difference between following along and missing half of it.
Lectures and conferences are where predictive captioning shines, because the conditions suit it: usually one person speaking at a time, a microphone, a topic the model can follow, and a setting where a small error here and there does not change anyone’s life. For that, AI captioning has become a legitimately useful access tool, and a free or low-cost one.
Where it still fails, and why that matters
The same prediction that smooths everyday speech becomes a liability the moment the stakes rise.
Medical settings. Drug names, dosages, anatomy, and diagnoses are exactly the rare, precise words a context model is most likely to “correct” into something more common and completely wrong. A system tuned to guess the likely word is dangerous when the right word is unusual and the difference is clinical.
Legal settings. Legal transcription needs specialized terminology and exact wording, and the meaning of a sentence can turn on a single word. In a courtroom, a deposition, or a tribunal, a confident-but-wrong caption is worse than a gap, because it reads as certain.
Multiple speakers and crosstalk. Accuracy drops with background noise, distance from the microphone, accents, and overlapping voices. A panel where people interrupt each other, a family meeting, a noisy clinic waiting room: these are the conditions where AI captioning degrades fastest, and they are common.
There is a quieter risk underneath all of this. When the system guesses, it guesses confidently. It does not flag the words it was unsure about. A human captioner who mishears something will often signal it; the AI just prints its best guess in the same clean font as everything it got right. For the reader, there is no visible difference between a word the system knew and a word it invented.
AI captioning is not the same as CART
This is the distinction that matters most, and it is one the disability community has been clear about. CART, Communication Access Realtime Translation, is live captioning produced by a trained human captioner. It remains the standard for formal accommodations precisely because a person can be held to an accuracy standard and can handle accents, crosstalk, and specialized vocabulary in ways the automated systems still cannot.
One note on the vocabulary. CART is the American term, and you will see it used in Canada too, but Google’s own accessibility researchers describe it as the US service, alongside palantypists in the United Kingdom and speech-to-text reporters elsewhere. The word travels; the funding and the legal rules do not. If you are in Canada, your right to an accommodation does not come from the Americans with Disabilities Act. It comes from the Canadian Human Rights Act or your provincial human rights code, depending on who you are dealing with, and those impose a duty to accommodate up to the point of undue hardship. The Accessible Canada Act adds barrier-removal duties on federally regulated organizations on top of that.
The quality worry is not new and it is not imported. The Canadian Association of the Deaf has said plainly that increasing reliance on voice-recognition technology has led to an increase in poor quality captioning, and that on-site human captioners give a far better chance of true synchronicity. That is a Canadian organization of Deaf people describing what has happened to captioning in this country as the automated tools moved in.
So when someone offers to swap automatic captions in for the human captioner you requested, you are not being difficult by saying no. An accommodation is a legal obligation, not a courtesy, and it is not satisfied by a cheaper tool that fails in the exact situations where accuracy is the point.
How to use it well
- Reach for AI captioning for everyday access: lectures, meetings, webinars, casual conversation in quiet rooms. It is fast, often free, and good enough where the cost of an occasional error is low.
- Do not rely on it alone for anything where a wrong word changes the outcome: medical appointments, legal proceedings, financial or contractual conversations. Ask for a human captioner, request a written follow-up, or have someone confirm the details.
- If you are entitled to CART as an accommodation, you can decline a substitution of automated captions. Accuracy that cannot be guaranteed is not equivalent access.
- Watch for the silent error. If a captioned sentence does not quite make sense, treat it as a possible miss rather than assuming you misunderstood. The system will not tell you when it guessed.
- In multi-speaker settings, expect more errors and plan for it: ask people to speak one at a time, get close to the microphone, and confirm anything important.
The technology has earned a real place in daily life, and that is worth celebrating without overselling it. The honest line is the boring one: predictive captioning is a strong everyday tool and a poor substitute for a trained human when precision is the whole point. Use it for what it is good at, and hold the line on real accommodation everywhere else.
Sources
- Radford et al., Robust Speech Recognition via Large-Scale Weak Supervision (the Whisper paper), OpenAI, December 2022
- Captioning and Video Accessibility, Canadian Association of the Deaf / Association des Sourds du Canada
- Real-time Continuous Transcription with Live Transcribe, Google Research
- Transcription and Captioning, Colgate University Accessibility Resources (American, describes US practice)
Related on Living Unlimited
- Connected by Touch: The Communication Tech DeafBlind Canadians Use, and What Is New
- Coding With a Disability: The Tools, the Workflows, the Real Constraints
- Air Travel Complaints That Actually Work: How to File and What to Expect
