Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms

Multimodal Transformer Brain Encoder Captures fMRI Responses during Naturalistic Video Viewing

Pinyuan Feng1, Hossein Adeli1, Richard Antonello1, Ethan Hwang1, Nikolaus Kriegeskorte1; 1Columbia University

Presenter: Pinyuan Feng

A primary goal of neuroscience is to understand how the brain integrates multimodal information under naturalistic settings. Here, we introduce Multimodal Transformer Brain Encoder (MTBEn), a neural encoding model for predicting parcel-level fMRI responses during naturalistic video viewing. MTBEn extends prior vision-based approaches by introducing cross-attention, in which learnable parcel-level queries attend to multimodal representations from visual, auditory, and linguistic inputs. This mechanism enables parcel-specific routing of multimodal representations that can be linearly mapped to fMRI responses. We evaluated MTBEn on in- and out-of-distribution datasets across four subjects, showing accurate neural response prediction under naturalistic viewing and generalization across stimulus distributions. These results support cross-attention–based multimodal integration as a useful computational framework for modeling brain responses to naturalistic stimuli.

Topic Area: Methods, Tools, Theory & Neural Coding