Keynotes | K&Ts | GACs | Talks | Posters | Search
Language & Audition
Contributed Talk Session: Thursday, August 6, 11:15 am – 12:15 pm, Skirball Theater
Invariant correlation: a general and straightforward method to quantify invariance in time-varying neural signals
Talk 1, 11:15 am
Jarrod M. Hicks1, Dana Boebinger1, Kirill V Nourski2, Matthew A. Howard2, Christopher M. Garcia2, Thomas Wychowski1, Webster Pilcher1, Samuel Victor Norman-Haignere1; 1University of Rochester, 2University of Iowa
Presenter: Jarrod M. Hicks
Recognizing information in natural stimuli is challenging because sensory inputs for different instances of the same feature often vary substantially. Successful recognition thus requires neural systems to generate representations that are invariant to such variation, posing a fundamental computational challenge for sensory coding. Much remains unknown about how sensory systems code invariant information in complex, temporally dynamic stimuli such as speech, in part due to methodological challenges in measuring and modeling invariance from noisy, time-varying neural responses. Here, we introduce the “invariant correlation”, a straightforward and general method for directly quantifying the strength and temporal dynamics of invariance from time-varying neural signals and computational models. We demonstrate its utility by quantifying invariance to acoustic variation in neural representations of phonemes using spatiotemporally precise intracranial recordings from human auditory cortex. Our method successfully revealed substantial diversity in the strength and dynamics of invariance across the human auditory cortex and across different phonetic features within individual electrodes. We illustrate how many of these empirically measured patterns can be explained by applying the same analysis to the predictions from a simple computational model based on spectrotemporal tuning in a cochleagram representation of sound. Although we focus here on speech, the invariant correlation is a broadly applicable tool for characterizing the organization and computational mechanisms underlying invariant representations of dynamic stimuli across sensory modalities.
Data Diversity Drives the Emergence of Symbolic Mechanisms in LLMs
Talk 2, 11:25 am
Melody Zixuan Li1, Taylor Whittington Webb2; 1McGill University, 2Université de Montréal
Presenter: Melody Zixuan Li
Recent work has identified an emergent three-stage symbolic architecture in large language models (LLMs) supporting abstract reasoning. What training conditions drive its emergence? We train transformers from scratch on abstract sequence tasks and test how data diversity shapes symbolic mechanisms. Using a combination of mechanistic interventions, and behavioral assessments of systematic (i.e., out-of-distribution) generalization, we find that both mechanistic and behavioral signatures of symbol processing scale with data diversity, suggesting that this may be a key factor driving the emergence of symbolic computation in LLMs.
Functional Localization as a Path to a Cumulative Research Enterprise in Human Intracranial Research
Talk 3, 11:35 am
Suseendrakumar Duraivel1, Eghbal A. Hosseini2, Colton Casto3, Rui Xu1, Ashley Walton4, Alan Bush4, Angelique Paulk4, Sydney Cash4, Mark Richardson4, Evelina Fedorenko1; 1Massachusetts Institute of Technology, 2Google DeepMind, 3Harvard University, 4Massachusetts General Hospital
Presenter: Suseendrakumar Duraivel
Meaningful comparison of neuroscience findings requires assurance that the results pertain to the same brain region across individuals and studies. Human intracranial research enables direct cortical recordings with millisecond resolution, but electrode locations are idiosyncratic across individuals and typically treated as belonging to the same functional system based on coarse anatomy. However, outside primary sensory areas, macroanatomy is an unreliable guide to function due to substantial inter-individual variability in functional topographies. Here, we advocate for functional localization in human intracranial research and demonstrate its feasibility using the language system as a test case. We show that (1) a localizer developed for fMRI reliably identifies language-responsive electrodes; (2) high-gamma power and BOLD signal strength are correlated within participants; and (3) a probabilistic fMRI atlas predicts whether an electrode is language-responsive better than anatomy alone. Thus, functional localization extends straightforwardly to intracranial research and can facilitate across-study comparisons.
Investigating Invariances in Auditory Event Categorization with Model Metamers
Talk 4, 11:45 am
Hee So Kim1, Elizabeth J. Lee1, Malinda McPherson-McNato2, Abigail Noyce1, Jenelle Feather1; 1Carnegie Mellon University, 2Purdue University
Presenter: Hee So Kim
Real-world acoustic inputs contain rich sensory information that we parse into discrete auditory objects and categories. Although deep neural networks (DNNs) are increasingly used to model auditory perception, the field lacks rigorous behavioral benchmarks, particularly for auditory event categorization. Here, we developed a 25-way categorization paradigm for broad classes of natural sounds to test whether categorical invariances of DNNs align with those of human observers. We first confirmed that humans could reliably categorize the natural sounds, demonstrating that our paradigm is well-suited for testing invariances in auditory categories. To probe model invariances, we evaluated human recognition of `model metamers' (synthetic stimuli matched to the model's internal activations for each natural stimulus) for a wide range of architectures trained on speech or auditory event recognition. Evaluating widely-used public models, we found that human recognition of auditory event model metamers was generally influenced by the training task and data distribution; speech models trained on standard, curated datasets produced less recognizable metamers than auditory event recognition models. We additionally analyzed a controlled set of models to directly investigate the influence of training task and adversarial training, revealing that improved metamer recognition induced by adversarial training is task-dependent. However, even in the best-performing models, we observed a sharp decline in human recognition at the final classification layer compared to the penultimate representation layer. Overall, our results suggest that while invariances in modern architectures better align with human observers for auditory event categorization, there is still a large discrepancy between the categorical invariances of auditory neural networks and the invariances of human observers. Code and models are available at https://github.com/Feather-Lab/env-sound-metamers.
Deep Learning Models of Attention Reveal Peripheral Encoding as a Bottleneck for Selective Listening Through Cochlear Implants
Talk 5, 11:55 am
Annesya Banerjee1, Ian M. Griffith1, Josh Mcdermott2; 1Harvard University, 2Massachusetts Institute of Technology
Presenter: Annesya Banerjee
Humans can selectively listen to target sounds in noisy environments. This ability is limited in users of cochlear implants (CIs), which attempt to restore hearing in deaf individuals by electrically stimulating the auditory nerve. To investigate the underlying causes, we optimized deep neural networks to recognize speech from a cued talker in multi-talker mixtures, using simulated binaural auditory nerve input from either a normal cochlea or a simulated cochlear implant. Attentional selection was implemented via feature-based gains, allowing models to selectively enhance target features while suppressing distractor features. Models with normal cochlear input successfully used both spatial and voice cues to recognize cued talkers while ignoring distractors, mirroring the performance of normal-hearing humans. By contrast, models with simulated CI stimulation performed notably worse, showing much less benefit from both voice and spatial cues. The presence of these limitations in models optimized to recognize speech from CI input suggests that attention deficits in CI users reflect limitations in the information available from current CI devices, underscoring the importance of developing alternative stimulation algorithms.
Lexical representations are sequenced through a micro-scale dynamic code in inferior frontal gyrus
Talk 6, 12:05 pm
Irmak Ergin1, Atlas Kazemian1, Ernesto Rojas1, Nick Hahn1, Foram Kamdar2, Christina Kim Vo1, Sasi Sathya Madugula1, Ryan Z Wang1, Chaofei Fan1, Erin Kunz1, Donald Avasino3, Leigh Hochberg4, Jaimie M. Henderson1, Francis R Willett1, Laura Gwilliams1; 1Stanford University, 2University of Illinois at Chicago, 3Massachusetts General Hospital, 4Brown University
Presenter: Irmak Ergin
Naturalistic speech is transient and strictly time-ordered, creating a computational challenge for comprehension: the brain must preserve and integrate earlier input while new words continue to arrive. One proposed algorithmic solution is that linguistic representations are maintained dynamically across changing neural populations over time, but it remains unclear how such an algorithm is implemented at the neuronal scale for higher-level linguistic representations. Using microelectrode recordings from the inferior frontal gyrus (IFG) during continuous speech listening, we show that lexical information is decodable for over 2 seconds, which overlaps with multiple words into the future. Neural activity evolves along state-space trajectories, which encode both lexical content and elapsed processing time. These trajectories preserve traces of earlier speech input as new input arrives, while also encoding relative sequence order. Together, these findings provide the first micro-scale evidence for how dynamic neural sequencing is implemented by IFG neurons during continuous speech comprehension.