Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster C in Poster Session C: Wednesday, August 5, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Human-aligned Universal Audio Representations with Contrastive-Equivariant Self-Supervised Learning

Ian M. Griffith1, Thomas Edward Yerxa2, Josh Mcdermott3, Jenelle Feather4; 1Harvard University, 2Meta, 3Massachusetts Institute of Technology, 4Carnegie Mellon University

Presenter: Ian M. Griffith

Human auditory experience consists of a wide range of environmental sounds in addition to music and speech. Yet, computational models of auditory processing are often trained for a single domain, with resulting representations that do not easily transfer to other tasks. Here, we adapt Contrastive-Equivariant SSL (CE-SSL) to the audio domain to learn a universal audio representation. The learning algorithm operates on superpositions and transformations of source signals, enforcing equivariance to concurrent source mixing and other acoustic transformations. We demonstrate the viability of the resulting models with three sets of results. First, the learned representations transfer to tasks in three different domains (word recognition, musical instrument identification, and environmental sound identification). Second, the learned representations preserve variance to non-categorical acoustic features, enabling superior zero-shot transfer to acoustic discrimination and paralinguistic tasks. Third, we show that the learned representations predict fMRI responses on-par with representations from supervised models, achieving strong neural predictivity without explicit labels. Our approach easily scales to large unlabeled datasets. The results demonstrate a promising self-supervised training method to achieve general purpose audio representations, and point the way towards human-aligned auditory models derived from unlabeled naturalistic datasets.

Topic Area: Auditory, Speech & Language Processing