Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms

Identifying Multimodal-Specific Regions with Naturalistic Stimuli Using a Contrastively-Trained Neural Network

Yixing Wang1, Faxin Zhou1, Adeen Flinker1; 1New York University

Presenter: Yixing Wang

The brain integrates auditory and visual information into a unified percept. However, most audiovisual neuroscience studies focus on low-level features or rely on separate unimodal representations, limiting their ability to capture fused content across modalities. To address this gap, we leverage ImageBind, a contrastively trained model that maps audio and visual inputs into a shared semantic space, together with rare intracranial electroencephalography (iEEG) signals from 19 neurosurgical participants during naturalistic movie viewing. By applying a multivariate temporal response function (mTRF) encoding framework, we characterize modality-specific brain networks for both audio and visual streams. Importantly, we find that combined audiovisual embeddings are more strongly encoded than both unimodal embeddings in a subset of electrodes, particularly in superior temporal regions, highlighting their role in semantic-level audiovisual integration. These results demonstrate that multimodal deep neural networks provide a promising framework for investigating audiovisual integration under naturalistic viewing.

Topic Area: Methods, Tools, Theory & Neural Coding