Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms

Emergent Eye Movements for Scene Understanding in a Model of Ventral and Dorsal Visual Processing

Alexander Kroner1, Mathis Pink2, Adrien Doerig3, Tim C Kietzmann1; 1Universität Osnabrück, 2MPI-SWS, 3Freie Universität Berlin

Presenter: Alexander Kroner

We continuously move our eyes to sample high-acuity information from the visual world around us. Computational models aiming to replicate this behavior are typically trained on eye-tracking data to mimic the fixation patterns of people looking at, for example, natural scenes. While such models achieve remarkable predictive performance, they do not address the question of what eye movements are made for and may therefore fail to capture the actual computations that give rise to them. Here, we introduce a normative model that predicts semantic caption embeddings by actively selecting where information is sampled from in a scene. The network comprises two processing streams, akin to the ventral and dorsal pathways in the visual system: one integrates glimpses over time to predict scene caption embeddings, while the other determines where to fixate next. Importantly, the glimpse selection policy is optimized solely for performance on the scene understanding task and receives no supervision from empirical eye-tracking data. We find that the emergent fixation strategy generalizes to novel scenes and resembles human fixation density maps on two datasets, under both free-viewing and scene-description conditions. This result suggests that predicting scene caption embeddings provides an ecologically motivated objective for fixation modeling, in which human-like eye movements emerge as an efficient sampling strategy rather than serve as the explicit target of optimization.

Topic Area: Computational Models of Vision & Visual Cortex