Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms
Affective Factors Predict Selective Attention in Vision-Language Models
Michael Zhu1, Yuan Chang Leong1; 1University of Chicago
Presenter: Michael Zhu
Emotionally salient stimuli capture human attention more effectively than neutral ones. Vision-language models trained on large-scale human-generated image–text pairs may acquire similar biases in what they preferentially encode. We tested this hypothesis by presenting six CLIP models with side-by-side images drawn from the OASIS affective image set, and quantified attentional bias using two complementary measures. Trial-level regressions revealed that arousal consistently predicted attentional bias across all models and metrics, such that more arousing images were prioritized in the models' representations. Valence effects were weaker and less consistent. These findings suggest that large-scale multimodal training may transmit not only semantic correspondence but also human-like affective biases in visual attention.
Topic Area: Computational Models of Vision & Visual Cortex