Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
Mirrored from arXiv — NLP / Computation & Language for archival readability. Support the source by reading on the original site.
Computer Science > Computation and Language
Title:Benchmarking Human and Automatic Speech Recognition of Diverse Speech: Initial Results
Abstract:Humans are often considered to be the best listeners and seen as the upper-bound performance of automatic speech recognition (ASR) systems. We present a preliminary comparison of the performances of state-of-the-art ASR systems and Dutch native listeners on the recognition of "diverse" speech, specifically Dutch child and older adults' speech and Flemish. Google Telephony outperformed the other ASR systems. Importantly, the ASR systems showed similar performance to the listeners, and in specific cases even outperformed them. Slight performance differences between the listeners and ASR systems were found related to speaker's age and regional accents and utterance length. Future research should focus on making ASR systems more robust to acoustic variability related to aging and regional accents. A comparison of ASR recognition performances on the test stimuli and the full Jasmin-CGN test sets showed the influence of the specific test sets on the conclusions regarding benchmarking human and ASR performance.
| Comments: | 7 pages, 4 figures |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.19049 [cs.CL] |
| (or arXiv:2607.19049v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.19049
arXiv-issued DOI via DataCite (pending registration)
|
Access Paper:
- View PDF
- HTML (experimental)
- TeX Source
References & Citations
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
More from arXiv — NLP / Computation & Language
-
Unified Hallucination Fuzzing for Multimodal Large Language Models
Aug 11
-
DocAtlas: Long-Document Understanding as Mutable-State Interaction
Aug 11
-
WuYuEval: A Multi-Level Benchmark for Large Language Models in Solid Waste Management
Aug 11
-
Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
Aug 11
Discussion (0)
Sign in to join the discussion. Free account, 30 seconds — email code or GitHub.
Sign in →No comments yet. Sign in and be the first to say something.