Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms
Visual and Auditory Deep Learning Models Capture Neural Representations of Naturalistic Social Interaction in the Superior Temporal Sulcus
Itai peleg1, Shiri Almog1, Idan Daniel Grosbard1, Maya Kadushin1, Nitzan Guy1, Ido Tavor1, Galit Yovel1; 1Tel Aviv University
Presenter: Itai peleg
The superior temporal sulcus (STS) selectively processes multimodal social information, yet modeling its responses has relied on specific categories (e.g., face, body, voice), or low-dimensional human annotations. We tested whether pre-trained vision-language (CLIP) and audio-language (CLAP) deep learning models better predict STS activity during naturalistic movie viewing. The joint CLIP-CLAP model significantly outperformed a social-affective annotation model across all ROIs. PCA on encoding model weights and projecting on time points revealed interpretable dimensions: a social interaction along the primary dimension for both visual and auditory models. A shot size (close up) gradient for CLIP, and a novel emotional prosody gradient for CLAP along the second dimension. These models reveal that the STS encodes interpretable dimensions of social information across visual and auditory modalities.
Topic Area: Methods, Tools, Theory & Neural Coding