Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms
Task-specific scene semantics from language predicts human gaze in navigation and visual search
Michelle R. Greene1, Bruce Hansen2; 1Barnard College, 2Colgate University
Presenter: Michelle R. Greene
Human gaze is profoundly task-driven, yet most computational models treat attention as primarily stimulus-driven. Here, we use natural language as a structured means of generating task-specific priority maps. We collected human scene descriptions under navigation and object identification goals, and trained convolutional neural networks to predict each task’s embeddings from images, using gradient-based attribution to extract spatial priority maps from the trained networks. We tested these maps against human gaze during real-world navigation, and visual search. The navigation map was the only model to achieve positive information gain over center bias during real-world navigation, outperforming all comparison models including DeepGaze IIE, which failed on mobile eye-tracking data despite strong performance on laboratory-based search. Critically, the relative predictive advantage of the two semantic maps reversed across tasks: navigation maps outperformed object maps during navigation, while object maps outperformed navigation maps during visual search. This double dissociation demonstrates that natural language about scenes encodes task-specific spatial priorities.
Topic Area: Computational Models of Vision & Visual Cortex