Key takeaways
- Speech input hit 153 words per minute in English versus 52 WPM for a touchscreen keyboard — 2.93× faster. Ruan et al., arXiv, 2016/2018
- Dictation was also more accurate: 5.30% corrected error rate for speech versus 11.22% for typing. Ruan et al., 2016/2018
- Even the fastest typists in a 136-million-keystroke dataset reach 120 WPM or more — still short of measured speech input. Dhakal et al., CHI 2018
- In Mandarin, speech reached 123 WPM against 43 WPM typed — 2.87× faster. Ruan et al., 2016/2018
- The best reported word error rate on the 2000 Switchboard set is 5.1%. Microsoft, 2017
- OpenAI's Whisper was trained on 680,000 hours of multilingual audio. Radford et al., 2022
- Across five commercial systems, average word error rate was 0.35 for black speakers versus 0.19 for white speakers. Koenecke et al., PNAS, 2020
- US medical transcriptionist jobs: 42,000 in 2025, projected −4% by 2035. US Bureau of Labor Statistics
- Court reporters and simultaneous captioners earn a median $72,420, with employment projected to change 0% to 2035. US Bureau of Labor Statistics
- Weekly meetings per Microsoft Teams user rose 153% from the start of the pandemic. Microsoft Work Trend Index, 2022
- 57% of meetings are ad hoc calls with no calendar invite. Microsoft Work Trend Index, 2025
- Workers are interrupted every 2 minutes — about 275 times a day. Microsoft Work Trend Index, 2025
- 60% of US adults read the AI-generated summary at the top of search results. Pew Research Center, June 2026
- 49% of US adults now use AI chatbots, up from 33% in 2024. Pew Research Center, June 2026
- 19.95% of EU enterprises used AI in 2025; 7.22% specifically used speech recognition. Eurostat
- One model now covers 1,107 languages for speech recognition. Pratap et al., 2023
- 430 million people need rehabilitation for disabling hearing loss today; 2.5 billion are projected to have some hearing loss by 2050. World Health Organization
- With ambient AI documentation, median time on notes fell from 7.1 to 6.1 minutes per appointment across 1,547 clinicians. JAMA Network Open, 2026
How much faster is speaking than typing?
Speech input was measured at 153 words per minute in English against 52 words per minute on a touchscreen keyboard, a 2.93× speed-up. In Mandarin the same study measured 123 WPM for speech against 43 WPM typed, a 2.87× speed-up.
The accuracy result is the one most people get wrong. Speech was not a trade-off against correctness: the corrected error rate was 5.30% for speech versus 11.22% for the keyboard. Uncorrected error rates ran the other way, 1.30% for speech against 0.79% for typing.
The comparison used Baidu's Deep Speech 2 against the built-in iOS keyboard on an iPhone 6 Plus, under laboratory conditions, for short messages.
For context on the typing side of that comparison, the largest study of everyday typing — 136 million keystrokes from 168,000 volunteers — found the fastest typists “reach speeds of even 120 WPM or more”, while the largest and slowest group of typists sat around 46 WPM. Formal training barely moves it: trained typists were only 5 WPM faster on average than untrained ones.
So the ceiling matters. Speech input at 153 WPM is not merely faster than average typing — it is faster than the fastest typists in a dataset of that size, without practice.
Source: Vivek Dhakal, Anna Feit, Per Ola Kristensson, Antti Oulasvirta, “Observations on Typing from 136 Million Keystrokes”, CHI 2018, Aalto University.
Source: Sherry Ruan, Jacob O. Wobbrock, Kenny Liou, Andrew Ng, James Landay, “Comparing Speech and Keyboard Text Entry for Short Messages in Two Languages on Touchscreen Phones”, arXiv:1608.07323, submitted 2016, revised 2018.
How accurate is speech recognition at its best?
The lowest widely cited benchmark figure is a 5.1% word error rate on the 2000 Switchboard evaluation set, reported by Microsoft in 2017. Note what that number is and is not: it is conversational telephone speech under benchmark conditions, not a guarantee for your meeting room. The paper itself makes no claim of parity with human transcribers.
Scale is the other half of the story. Whisper was trained on 680,000 hours of multilingual and multitask supervision — the shift from per-user voice training to models that generalise from very large, messy datasets.
Sources: W. Xiong, L. Wu, F. Alleva, J. Droppo, X. Huang, A. Stolcke (Microsoft), “The Microsoft 2017 Conversational Speech Recognition System”, arXiv:1708.06073, 2017. A. Radford, J. W. Kim, T. Xu, G. Brockman, C. McLeavey, I. Sutskever (OpenAI), “Robust Speech Recognition via Large-Scale Weak Supervision”, arXiv:2212.04356, 2022.
Is speech recognition equally accurate for everyone?
No. Across five commercial systems from Amazon, Apple, Google, IBM and Microsoft, the average word error rate was 0.35 for black speakers against 0.19 for white speakers — nearly double the error rate for the same task.
The study used structured interviews with 42 white speakers and 73 black speakers across five US cities, totalling 19.8 hours of audio. The gap persisted when both groups spoke identical phrases, which the authors attribute to the underlying acoustic models rather than vocabulary or grammar.
This is the single most important number on this page for anyone shipping voice software: benchmark accuracy and accuracy for your actual users are different quantities.
Source: Allison Koenecke, Andrew Nam, Emily Lake, Joe Nudell, Minnie Quartey, Zion Mengesha, Connor Toups, John R. Rickford, Dan Jurafsky, Sharad Goel, “Racial disparities in automated speech recognition”, Proceedings of the National Academy of Sciences 117(14):7684–7689, 7 April 2020. doi:10.1073/pnas.1915768117 (open record: PubMed)
What is automation doing to transcription work?
It is shrinking the jobs that only transcribe, while leaving the jobs that certify or interpret largely intact. US Bureau of Labor Statistics projections for 2025–2035:
| Occupation | Jobs (2025) | Median pay (May 2025) | Projected 2025–35 |
|---|---|---|---|
| Medical transcriptionists | 42,000 | $40,410 ($19.43/hr) | −4% (−1,900 jobs) |
| Court reporters & simultaneous captioners | 19,900 | $72,420 | 0% (no change) |
| Interpreters & translators | 73,900 | $60,170 | +2% (+1,500 jobs) |
Court reporting is the instructive case: employment is flat, yet about 1,800 openings are projected each year over the decade, driven by retirement rather than growth.
Source: US Bureau of Labor Statistics, Occupational Outlook Handbook: Medical Transcriptionists, Court Reporters and Simultaneous Captioners, Interpreters and Translators. Pages last modified 27 August 2026.
How much of the working week goes into meetings?
Meeting load grew sharply and then changed shape. Microsoft measured a 153% increase in weekly meetings per Teams user from the start of the pandemic, and a 46% increase in overlapping (double-booked) meetings per person year over year.
Its later telemetry describes a day that is fragmented rather than merely full:
- 57% of meetings are ad hoc calls with no calendar invite.
- 50% of all meetings fall in two windows: 9–11am and 1–3pm.
- Tuesday carries the heaviest load at 23% of meetings; Friday the lightest at 16%.
- 1 in 10 scheduled meetings is booked at the last minute.
- Nearly a third of meetings span multiple time zones, up 35% since 2021.
- Meetings after 8pm are up 16% year over year.
- Large meetings of 65+ attendees are the fastest-growing type.
- Workers are interrupted every 2 minutes, roughly 275 times a day.
That pattern — unscheduled, overlapping, out-of-hours — is why note-taking that depends on someone being free to type has stopped working. See how bot-free meeting capture works.
Sources: Microsoft, “Hybrid Work Is Just Work. Are We Doing It Wrong?”, Work Trend Index, 22 September 2022. Microsoft, “Breaking down the infinite workday”, Work Trend Index special report, 17 June 2025.
Does AI note-taking actually save measurable time?
Yes, and the best-measured evidence comes from clinical documentation. Across 1,547 clinicians and 16,149 observations, median time spent on notes fell from 7.1 minutes to 6.1 minutes per appointment after ambient AI documentation was introduced.
The modelled effects were smaller but consistent: a 0.26 minute decrease per note in the first month, a sustained 0.38 minute additional monthly decrease in after-hours documentation, and an immediate increase of 7.40 relative value units per month. The study found no significant change in the clinician efficiency profile score.
One caveat worth repeating, because roundups routinely drop it: this study measured documentation time and productivity, not burnout. Figures attributing burnout reductions to this paper are not in it.
Source: “Ambient Artificial Intelligence Use and Clinician Documentation Burden, Productivity, and Efficiency”, JAMA Network Open, 29 May 2026. jamanetwork.com (open access: PubMed Central)
What rules apply to AI in hiring interviews?
In New York City, employers using an automated employment decision tool must have had it bias-audited within the previous year, must publish a summary of the audit results, and must notify candidates at least 10 business days before use. The Department of Consumer and Worker Protection began enforcing Local Law 144 of 2021 on 5 July 2023.
If you are on the candidate side of that, our guide to AI interview assistants covers what the tools do and where the lines are.
Source: New York City Department of Consumer and Worker Protection, Automated Employment Decision Tools, Local Law 144 of 2021.
How many people and companies actually use this?
About half the US adult population now uses AI chatbots: 49% of US adults in early 2026, up from 33% in 2024, with ChatGPT itself at 18% back in 2023. Platform shares run ChatGPT 44%, Gemini 24%, Copilot 17%, Meta AI 14%, Grok 8%, Claude 6% and Character.ai 3%.
The number with the widest consequences for anyone publishing online: 60% of US adults say they read the AI-generated summaries at the top of search results. Being the source those summaries draw from is now a distribution channel in its own right.
Source: Pew Research Center, “Americans and AI 2026: Chatbots, Smart Devices and Views on Impact”, 17 June 2026; survey of 5,119 US adults, fielded 17–23 February 2026.
On the organisational side, 88% of surveyed organisations use AI in at least one capacity and 70% use generative AI in at least one business function. Generative AI reached 53% adoption in three years, faster than personal computers or the internet.
Source: Stanford HAI, 2026 AI Index Report, Economy chapter.
European figures come from a statistical agency rather than a vendor survey, which makes them the better citation. In 2025, 19.95% of EU enterprises used AI technologies, up 6.47 percentage points on 2024. By technology: text mining 11.75%, generative media 9.55%, natural language generation and speech synthesis 8.76%, and speech recognition 7.22%.
Adoption is wildly uneven. Denmark 42.03%, Finland 37.82% and Sweden 35.04% lead; Romania 5.21%, Poland 8.36% and Bulgaria 8.55% trail. By size, 55.03% of large enterprises use AI against 30.36% of medium and 17% of small ones.
Source: Eurostat, “Use of artificial intelligence in enterprises”, Statistics Explained, 2025 reference year.
How many languages can speech recognition handle?
A single multilingual model now covers 1,107 languages for speech recognition, with speech synthesis for the same number. On the FLEURS benchmark it “more than halves the word error rate of Whisper on 54 languages”.
That matters beyond bragging rights: most of the world's languages have never had usable dictation at all, and the constraint has always been labelled data rather than model design.
Source: Vineel Pratap et al. (Meta AI), “Scaling Speech Technology to 1,000+ Languages”, arXiv:2305.13516, 2023.
The open-data side is smaller than people assume. Mozilla's Common Voice reported 2,500 hours of collected audio across 38 languages being collected (29 in the then-current release) from over 50,000 contributors, and an average character error rate improvement of 5.99 ± 5.48 across twelve target languages via transfer learning.
Source: Rosana Ardila et al. (Mozilla), “Common Voice: A Massively-Multilingual Speech Corpus”, arXiv:1912.06670, 2019, revised 2020. Figures are as published then; the corpus has grown since.
Who needs transcription and captions most?
Over 5% of the world's population — 430 million people, including 34 million children — need rehabilitation for disabling hearing loss today. Around 95.1 million children aged 5–19 live with hearing loss. By 2050, nearly 2.5 billion people are projected to have some degree of hearing loss and more than 700 million will require rehabilitation. Unaddressed hearing loss carries an annual global cost of almost US$1 trillion.
Source: World Health Organization, “Deafness and hearing loss” fact sheet, last updated 3 March 2026.
Captions are also becoming a legal obligation rather than a courtesy. The US Department of Justice's 2024 final rule requires state and local government websites and mobile apps to meet WCAG 2.1 Level AA — which includes captions for video — by 26 April 2027 for entities serving populations of 50,000 or more, and 26 April 2028 for smaller entities and special districts.
Source: US Department of Justice, “Nondiscrimination on the Basis of Disability: Accessibility of Web Information and Services of State and Local Governments”, ADA Title II final rule, published 24 April 2024.
How to cite this page
Every statistic above belongs to the organisation named beside it — cite them, not us. If you want to reference the collection itself:
ChadFlow, “Voice AI statistics 2026”, updated 12 September 2026. https://chadflow.app/voice-ai-statistics
If you find a figure here that has drifted from its source, or a source that has moved, tell us at [email protected] and we will correct it.
Frequently asked questions
How much faster is speaking than typing?
In a controlled study of short-message entry on an iPhone, speech input reached 153 words per minute in English against 52 words per minute for the touchscreen keyboard — 2.93 times faster (Ruan et al., arXiv, 2016/2018).
What is the lowest word error rate on the Switchboard benchmark?
Microsoft reported a 5.1% word error rate on the 2000 Switchboard evaluation set in its 2017 conversational speech recognition system paper.
Is speech recognition equally accurate for everyone?
No. Across five commercial systems, average word error rate was 0.35 for black speakers against 0.19 for white speakers (Koenecke et al., PNAS, 2020).
How many medical transcriptionist jobs are left in the US?
42,000 in 2025, projected to fall 4% (a loss of 1,900 jobs) between 2025 and 2035, according to the US Bureau of Labor Statistics Occupational Outlook Handbook.
Does AI note-taking actually save time?
In a study of 1,547 clinicians, median time spent on notes per appointment fell from 7.1 minutes to 6.1 minutes after ambient AI documentation was introduced (JAMA Network Open, 29 May 2026).
Put the 2.93× to work.
ChadFlow does dictation, meeting transcription with speaker labels, and live answers — on Mac, Windows and iPhone.
Download ChadFlowMethod: Each figure was checked against the source page or paper on 12 September 2026, and only included where the number appeared in the source itself. Vendor market-size projections were deliberately excluded: the published figures we could reach were not verifiable against a primary methodology. Where a widely repeated claim was not present in the paper it is attributed to — human-parity transcription, burnout reductions from the JAMA study — we say so rather than repeat it.