Keynotes | K&Ts | GACs | Talks | Posters | Search

Contributed Talk Session: Tuesday, August 4, 2:00 – 3:00 pm, Skirball Theater
Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Integrating vision and language in long-term memory metamers

Abe Leite1, Ritik Raina1, Alexandros Graikos2, Seoyoung Ahn3, Gregory J. Zelinsky1; 1State University of New York at Stony Brook, 2Stony Brook University, 3Hankuk University of Foreign Studies

Presenter: Abe Leite

Long-term memory (LTM) incorporates both visuo-spatial and semantic information. Recent work has studied the influence of visual information on short-term memory using a ‘metamerism’ paradigm: asking viewers to judge whether a pair of scenes separated by a delay are the same, when the second may have been computationally generated based on the first. That work’s generative model, MetamerGen, supports visual but not semantic conditioning, limiting its ability to study LTM phenomena. In this work, we present Seen2Scene-v2, an extended MetamerGen model incorporating both visual and text conditioning. Keeping visual input constant, we evaluate the impact of Seen2Scene-v2’s linguistic conditioning on recall rates in LTM. In session 1, participants viewed and described 150 images from MetamerGen’s subset of the Visual Genome dataset; in session 2, they indicated whether they recalled having seen each of 180 briefly-shown images – 30 unseen, 30 original, and 120 generated based on images from session 1. Our critical manipulation was the textual input to Seen2Scene-v2: the participant’s own description, another participant’s description of the same image (cross-viewer condition), or a description of a different image (random baseline). We found an early own-description benefit that faded over 3 weeks, and a persistent random-condition penalty. This indicates that personalized semantics exist in LTM, but decay over time, whereas core meaning lasts. In future work, we will test how visual input affects recall to compare the strength of visual and linguistic information in LTM.

Topic Area: Memory, Learning & Knowledge Structures