Effects of Talker Dialect, Gender & Race on Accuracy of Bing Speech and YouTube Automatic Captions

Rachael Tatman; Conner Kasten

2017 INTERSPEECH INTERSPEECH 2017

Effects of Talker Dialect, Gender & Race on Accuracy of Bing Speech and YouTube Automatic Captions

Abstract

This project compares the accuracy of two automatic speech recognition (ASR) systems — Bing Speech and YouTube’s automatic captions — across gender, race and four dialects of American English. The dialects included were chosen for their acoustic dissimilarity. Bing Speech had differences in word error rate (WER) between dialects and ethnicities, but they were not statistically reliable. YouTube’s automatic captions, however, did have statistically different WERs between dialects and races. The lowest average error rates were for General American and white talkers, respectively. Neither system had a reliably different WER between genders, which had been previously reported for YouTube’s automatic captions [1]. However, the higher error rate non-white talkers is worrying, as it may reduce the utility of these systems for talkers of color.

🧭 Keyword Pioneer — dialect variation

🐣 Hot Topic Early Bird — word error rate

🐝 Cross-Pollinator — Artificial Intelligence, Computer Science, Computer Vision, Data Science & Analytics, Deep Learning, Interdisciplinary, Machine Learning, Natural Language Processing, Speech & Audio

🌉 Interdisciplinary Bridge — Machine Learning and Speech & Audio

Authors

Rachael Tatman , Conner Kasten

Topics

Speech & Audio > Recognition > Automatic Speech Recognition Machine Learning > Learning Types > Fairness

Keywords

automatic speech recognition word error rate dialect variation gender effect racial effect automatic caption speech recognition accuracy speech recognition bia

Download PDF

Related papers

Description of the Munich-Passau Snore Sound Corpus (MPSSC) 2017

A Study on Replay Attack and Anti-Spoofing for Automatic Speaker Verification 2017

Binaural Reverberant Speech Separation Based on Deep Neural Networks 2017

Building Audio-Visual Phonetically Annotated Arabic Corpus for Expressive Text to Speech 2017

A Comparison of Danish Listeners’ Processing Cost in Judging the Truth Value of Norwegian, Swedish, and English Sentences 2017