Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

Predicting Sound Detection in Auditory Scenes Using Task-Optimized Neural Networks

Sagarika Alavilli1, Lakshmi Narasimhan Govindarajan2, Josh Mcdermott2; 1Harvard University, 2Massachusetts Institute of Technology

Presenter: Sagarika Alavilli

Detecting sounds that occur in the world is a fundamental task for our auditory system. Much is known about what limits detection of sounds in stationary noise. However, the factors that constrain detection of sounds in everyday scenes remain poorly understood. The combinatorial complexity of natural scenes makes exhaustive behavioral investigation impractical, motivating the use of computational models to generate hypotheses and targeted predictions. Here, we ask whether artificial neural network (ANN) representations optimized for environmental sound recognition predict human detection of individual sources within multi-source scenes. We randomly generated scenes in which the ANN model and a baseline spectrotemporal filter model made divergent predictions about which sources were most and least detectable, and tested these predictions against human behavioral performance. Human listeners showed the highest detection accuracy for sources predicted to be most detectable by the ANN model and the lowest accuracy for sources it predicted to be least detectable. By contrast, the spectrotemporal baseline model was not predictive of human performance. These findings suggest that representations optimized for sound recognition capture information relevant to human sound detection that is not available in traditional auditory models, and point to recognition-optimized models as a promising framework for studying auditory detection and salience.

Topic Area: Auditory, Speech & Language Processing