Keynotes | K&Ts | GACs | Talks | Posters | Search
Contributed Talk Session: Thursday, August 6, 11:15 am – 12:15 pm, Skirball Theater
Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms
Investigating Invariances in Auditory Event Categorization with Model Metamers
Hee So Kim1, Elizabeth J. Lee1, Malinda McPherson-McNato2, Abigail Noyce1, Jenelle Feather1; 1Carnegie Mellon University, 2Purdue University
Presenter: Hee So Kim
Real-world acoustic inputs contain rich sensory information that we parse into discrete auditory objects and categories. Although deep neural networks (DNNs) are increasingly used to model auditory perception, the field lacks rigorous behavioral benchmarks, particularly for auditory event categorization. Here, we developed a 25-way categorization paradigm for broad classes of natural sounds to test whether categorical invariances of DNNs align with those of human observers. We first confirmed that humans could reliably categorize the natural sounds, demonstrating that our paradigm is well-suited for testing invariances in auditory categories. To probe model invariances, we evaluated human recognition of `model metamers' (synthetic stimuli matched to the model's internal activations for each natural stimulus) for a wide range of architectures trained on speech or auditory event recognition. Evaluating widely-used public models, we found that human recognition of auditory event model metamers was generally influenced by the training task and data distribution; speech models trained on standard, curated datasets produced less recognizable metamers than auditory event recognition models. We additionally analyzed a controlled set of models to directly investigate the influence of training task and adversarial training, revealing that improved metamer recognition induced by adversarial training is task-dependent. However, even in the best-performing models, we observed a sharp decline in human recognition at the final classification layer compared to the penultimate representation layer. Overall, our results suggest that while invariances in modern architectures better align with human observers for auditory event categorization, there is still a large discrepancy between the categorical invariances of auditory neural networks and the invariances of human observers. Code and models are available at https://github.com/Feather-Lab/env-sound-metamers.
Topic Area: Auditory, Speech & Language Processing