Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms

Aligning the Brain with Language Models Through a Nonlinear and Multimodal Approach

Danny Dongyeop Han1, Yunju Cho1, Jiook Cha1, Jay-Yoon Lee1; 1Seoul National University

Presenter: Danny Dongyeop Han

Speech encoding models are widely used to test whether stimulus representations predict neural activity during naturalistic language comprehension. Recent work has shown that large language and speech models provide strong features for this task, yet most speech encoding studies still rely on linear mappings from unimodal features. This is restrictive as speech comprehension depends on distributed interactions between acoustic and linguistic information that are unlikely to be purely additive or linear. Here, we test whether a simple nonlinear multimodal encoder improves prediction of fMRI responses during naturalistic listening. We combine text and audio representations from LLaMA and Whisper, and compare linear and nonlinear encoders under matched settings. The best model, a multimodal MLP, improves average voxelwise r² by 17.2% and normalized correlation by 17.9% over the standard text-linear baseline. Matched control models show that these gains are not explained by dimensionality reduction or parameter count alone, but reflect the benefit of nonlinear multimodal fusion. Multimodal models improve prediction not only in auditory cortex but also in motor and somatosensory regions, with broader cortical gains amplified by nonlinearity. These findings suggest that linear unimodal practice leaves structured variance unexplained and that compact nonlinear multimodal encoders are a promising direction for naturalistic speech neuroscience.

Topic Area: Auditory, Speech & Language Processing