Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms
Modern Machine Systems Exhibit Superhuman Speech Recognition
Annika Magaro1, Gasser Elbanna1, Josh Mcdermott2; 1Harvard University, 2Massachusetts Institute of Technology
Presenter: Annika Magaro
Advances in machine learning have yielded machine systems that often approach human performance levels on real-world perception tasks. For speech recognition, it remains unclear whether the behavior of such models resembles that of humans. We addressed this question by measuring word recognition abilities in humans and models across a large set of speech distortions. Three recent industry speech recognition models trained on hundreds of years of speech exhibited high correlations with the pattern of human performance. However, in absolute terms they were superhuman on most types of distortion. Overall, the results show that superhuman performance is now possible with systems that are optimized over enough data, and that industry systems fall short as models of human perception as a result.
Topic Area: Auditory, Speech & Language Processing