Keynotes | K&Ts | GACs | Talks | Posters | Search

Contributed Talk Session: Thursday, August 6, 11:15 am – 12:15 pm, Skirball Theater
Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Deep Learning Models of Attention Reveal Peripheral Encoding as a Bottleneck for Selective Listening Through Cochlear Implants

Annesya Banerjee1, Ian M. Griffith1, Josh Mcdermott2; 1Harvard University, 2Massachusetts Institute of Technology

Presenter: Annesya Banerjee

Humans can selectively listen to target sounds in noisy environments. This ability is limited in users of cochlear implants (CIs), which attempt to restore hearing in deaf individuals by electrically stimulating the auditory nerve. To investigate the underlying causes, we optimized deep neural networks to recognize speech from a cued talker in multi-talker mixtures, using simulated binaural auditory nerve input from either a normal cochlea or a simulated cochlear implant. Attentional selection was implemented via feature-based gains, allowing models to selectively enhance target features while suppressing distractor features. Models with normal cochlear input successfully used both spatial and voice cues to recognize cued talkers while ignoring distractors, mirroring the performance of normal-hearing humans. By contrast, models with simulated CI stimulation performed notably worse, showing much less benefit from both voice and spatial cues. The presence of these limitations in models optimized to recognize speech from CI input suggests that attention deficits in CI users reflect limitations in the information available from current CI devices, underscoring the importance of developing alternative stimulation algorithms.

Topic Area: Auditory, Speech & Language Processing