Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster C in Poster Session C: Wednesday, August 5, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms
Eyes of the beholder: attention-guided metamers reveal divergent scene representations
Ritik Raina1, Abe Leite1, Alexandros Graikos2, Seoyoung Ahn3, Gregory J. Zelinsky1; 1State University of New York at Stony Brook, 2Stony Brook University, 3Hankuk University of Foreign Studies
Presenter: Ritik Raina
The human representation of visual scenes integrates high-resolution information sampled by viewing fixations with lower-resolution context extracted from the visual periphery. In this paper, we ask how individualized scene viewing by different people translates into individualized scene understanding. We do this by showing that scene metamers—generated scenes judged to be the same as originally viewed scenes—depend on the specific scanpath of fixations guiding the generation. To enable attention-guided scene generation we introduce Seen2Scene-v2, a latent diffusion model that generates images from scene-viewing fixations and blurred peripheral pixels. Human participants had to view a scene and then make a “same” or “different” judgment about a briefly presented second scene, many of which were fixation-guided generations from Seen2Scene-v2. The generative model was therefore integrated into our behavioral data collection. We obtained metamerism rates (the proportions of trials where generated images were confused with originals) under three attention conditions: the viewer's own fixations (own condition), another viewer's fixations on the same image (cross-viewer condition), and another viewer's fixations on an unrelated image (random condition). We found a consistent ordering of results: scenes generated using a participant's own attention were most likely to be judged as ”same”, with cross-viewer metamer rates being significantly lower and random-condition rates lower still. To explain these mistaken “same” judgments, we computed multiple representational similarity measures. Feature similarity across all levels of the visual hierarchy predicts metamerism, but the three-way separation between conditions persists when controlling for any one feature. We also compared our model to the current state-of-the-art in gaze-contingent scene generation and found that Seen2Scene-v2 better predicts our stratification of results. Based on our obtained differences between own- and cross-viewer conditions, despite being the same images, we conclude that people form different scene representations that are specific to their allocation of attention during viewing.
Topic Area: Computational Models of Vision & Visual Cortex