Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms
Do Better Visual Question Answering Models Attend More Like Humans?
Aiqing Li1; 1University of Arizona
Presenter: Aiqing Li
We examine whether improvements in visual question answering (VQA) models are associated with more human-like attention behavior. Using the VQA-HAT dataset, which provides human attention annotations, we analyze the attention patterns of a state-of-the-art model (ViLT-B/32) and compare them with both human attention and prior VQA models. While the model achieves strong task performance, its attention alignment with human annotations remains substantially below human-human agreement and comparable to earlier architectures. These observations suggest that improved model performance does not necessarily correspond to improved alignment with human attention. While the present analysis focuses on a single state-of-the-art architecture, the results provide preliminary evidence that predictive accuracy and human attention alignment may not improve in parallel.
Topic Area: Computational Models of Vision & Visual Cortex