Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

Visual and Auditory Deep Learning Models Capture Neural Representations of Naturalistic Social Interaction in the Superior Temporal Sulcus

Itai peleg1, Shiri Almog1, Idan Daniel Grosbard1, Maya Kadushin1, Nitzan Guy1, Ido Tavor1, Galit Yovel1; 1Tel Aviv University

Presenter: Itai peleg

The superior temporal sulcus (STS) selectively processes multimodal social information, yet modeling its responses has relied on specific categories (e.g., face, body, voice), or low-dimensional human annotations. We tested whether pre-trained vision-language (CLIP) and audio-language (CLAP) deep learning models better predict STS activity during naturalistic movie viewing. The joint CLIP-CLAP model significantly outperformed a social-affective annotation model across all ROIs. PCA on encoding model weights and projecting on time points revealed interpretable dimensions: a social interaction along the primary dimension for both visual and auditory models. A shot size (close up) gradient for CLIP, and a novel emotional prosody gradient for CLAP along the second dimension. These models reveal that the STS encodes interpretable dimensions of social information across visual and auditory modalities.

Topic Area: Methods, Tools, Theory & Neural Coding