Keynotes | K&Ts | GACs | Talks | Posters | Search
Learning & Memory
Contributed Talk Session: Tuesday, August 4, 2:00 – 3:00 pm, Skirball Theater
What Makes a Category Memorable? Human Conceptual Diversity Outpredicts Perceptual Similarity and Deep Neural Networks
Talk 1, 2:00 pm
Dyllan Simpson1, Mariagrazia De gioia2, Bria Lorelle Long1, Timothy Brady1; 1University of California, San Diego, 2University of Bologna
Presenter: Dyllan Simpson
Conceptual knowledge structures human memory, with better memory performance for more conceptual diverse information even when it is perceptually similar. To what extent do modern deep neural networks capture the conceptual structure relevant for memory performance? We investigated how human judgments of within-category diversity relate to memorability and to the representational structure of deep neural networks. Participants sorted exemplar images from 498 object categories (from the THINGS database) based on either conceptual similarity (function, purpose, typical context) or perceptual similarity (shape, color, texture). We measured category-level memorability using a recognition memory task with within-category foils. We then tested how memory performance related to human diversity scores and to diversity computed from CLIP, DINOv3, and VGG16 embeddings across multiple layers. Human conceptual and perceptual diversity were correlated (r = .46) but dissociable, with some categories showing high conceptual but low perceptual diversity (e.g., electrical plugs that look similar but function in different countries). Critically, conceptual diversity was the strongest predictor of memorability (R² = .57), and adding perceptual diversity or model-based measures did not improve prediction. Across models, correlations with human conceptual and perceptual judgments both increased with layer depth, and all three architectures converged to similarly good predictions at their final layers. However, none matched the predictive power of human conceptual diversity ratings in predicting human memorability. These findings extend prior work linking conceptual distinctiveness to memory by demonstrating that purely local, within-category structure, independent of a category's position in global semantic space, predicts how well individual exemplars are remembered. The gap between human and model-derived diversity estimates suggests that the typical implementations of vision models lack the flexible, context-sensitive feature weighting that humans deploy when organizing objects conceptually, highlighting an important constraint for computational accounts of human visual memory.
Spatial Structure Facilitates Category Learning through Structured Representations
Talk 2, 2:10 pm
Michelle B. Hefner1, Maria Ruz2, Christopher Summerfield3; 1Universidad de Granada, 2University of Granada, 3UK AI Security Institute
Presenter: Michelle B. Hefner
The spatial organization in which objects are learned may influence how they are later categorized. Participants first learned object-location associations in grid configurations with either a structured one-dimensional (1D) layout or an Interdigitated (pseudorandom) layout, and then categorized the same objects in a subsequent task containing no spatial information. Categorization was more accurate following learning in the structured 1D configuration, suggesting that spatial structure facilitates subsequent category learning. To account for this effect, we trained neural networks to learn object embeddings during the object-location association task and subsequently use those embeddings for categorization. In a rich learning regime (small embedding weights), representations preserved spatial structure, resulting in lower loss for 1D relative to Interdigitated conditions. In contrast, a lazy regime (large embedding weights) produced less structured representations and no difference between conditions. These results suggest that spatial structure shapes the geometry of learned representations, facilitating category learning. This framework generates testable predictions for ongoing fMRI work examining whether similar structure is reflected in neural representations during learning.
Looking at nothing during deliberation reflects optimal sampling from episodic memory
Talk 3, 2:20 pm
Jonathan Nicholas1, Sixing Chen1, Marcelo G Mattar1; 1New York University
Presenter: Jonathan Nicholas
Episodic memory supports deliberation by allowing our choices to be guided by past experiences, yet the dynamics of retrieval during decision making remain poorly understood. Here, we address this gap by using eye tracking to trace episodic retrieval during value-based choice. In the absence of visual information, participants directed their gaze toward the encoding locations of retrieved episodes, and these fixation patterns aligned with their decisions. We further formalized optimal retrieval in this setting and found that people's fixations are consistent with optimal sampling from episodic memory. These results demonstrate that eye gaze can serve as a window into episodic sampling and suggest that people optimally retrieve episodes to guide their choices.
Integrating vision and language in long-term memory metamers
Talk 4, 2:30 pm
Abe Leite1, Ritik Raina1, Alexandros Graikos2, Seoyoung Ahn3, Gregory J. Zelinsky1; 1State University of New York at Stony Brook, 2Stony Brook University, 3Hankuk University of Foreign Studies
Presenter: Abe Leite
Long-term memory (LTM) incorporates both visuo-spatial and semantic information. Recent work has studied the influence of visual information on short-term memory using a ‘metamerism’ paradigm: asking viewers to judge whether a pair of scenes separated by a delay are the same, when the second may have been computationally generated based on the first. That work’s generative model, MetamerGen, supports visual but not semantic conditioning, limiting its ability to study LTM phenomena. In this work, we present Seen2Scene-v2, an extended MetamerGen model incorporating both visual and text conditioning. Keeping visual input constant, we evaluate the impact of Seen2Scene-v2’s linguistic conditioning on recall rates in LTM. In session 1, participants viewed and described 150 images from MetamerGen’s subset of the Visual Genome dataset; in session 2, they indicated whether they recalled having seen each of 180 briefly-shown images – 30 unseen, 30 original, and 120 generated based on images from session 1. Our critical manipulation was the textual input to Seen2Scene-v2: the participant’s own description, another participant’s description of the same image (cross-viewer condition), or a description of a different image (random baseline). We found an early own-description benefit that faded over 3 weeks, and a persistent random-condition penalty. This indicates that personalized semantics exist in LTM, but decay over time, whereas core meaning lasts. In future work, we will test how visual input affects recall to compare the strength of visual and linguistic information in LTM.
Decomposing iconic memory decay with a joint-feature model of structured errors
Talk 5, 2:40 pm
Gal Vishne1, Zoe Haynes1, Miranda Ye2, Nicholas Turk-Browne3, Michael N. Shadlen1,4, John Morrison2; 1Columbia University, 2Barnard College, 3Yale University, 4Howard Hughes Medical Institute
Presenter: Gal Vishne
Iconic memory provides a brief, high-capacity store of visual information, yet its representational structure and temporal dynamics remain poorly understood. We address this question using a continuous-report paradigm enabling precise characterization of errors over time. Subjects briefly viewed arrays of colored items and, after a variable delay, reported the location associated with a probed color. Mixture models applied to color and location errors, revealed increasing guess rates over time and, in some cases, decreasing precision. However, these models fail to account for structured errors arising from feature interactions. To address this, we developed a joint-feature probabilistic model in which items are encoded with uncertainty in both color and location. Within this framework, errors commonly interpreted as swaps emerge naturally from uncertainty in the underlying representation. The joint model provided a better account of behavior across all subjects and delays and revealed that iconic memory decay reflects two distinct mechanisms: decreasing precision and increasing probability of retrieval failure. These findings constrain the format of iconic memory representations and provide a computational framework for characterizing memory performance over time.
Memory Subspaces for Temporal Locations Across Timescales in the Monkey Precuneus
Talk 6, 2:50 pm
Zhiyong Jin1, Ning Su2, Aakash Sarkar3, Xufeng Zhou1, Jiayu Cheng4, Makoto Kusunoki5, Sze Chai Kwok1; 1Duke Kunshan University, 2East China Normal University, 3University of California, San Francisco, 4Wuhan University, 5University of Oxford
Presenter: Sze Chai Kwok
Understanding how the brain organizes the temporal structure of experience across timescales is central to episodic memory. Here, we examine neural population dynamics in primate precuneus during a temporal order judgment task spanning seconds to minutes, and cross-day retrieval of intertwined episodes. We recorded large-scale population activity (∼3,000 neurons) across 3 days and analyzed it using generalized linear models and low-dimensional manifold methods. Single neurons showed consistent positive modulation for temporal locations (TLs) across both 2-s and 1-min delay conditions during retrieval. At the population level, neural activity formed structured low-dimensional manifolds in dPCA space, maintaining a subspace that persistently encoded the TLs of extracted frames throughout the TOJ period. Correct trials exhibited coherent, synchronized dynamics across different TLs, while behavioral errors were associated with manifold collapse and loss of dynamical synchrony. During cross-day retrieval, episodic states corresponding to the TL of clips viewed across both days rapidly separated into day-specific clusters within a shared state space, an effect absent in control conditions. The chronologically recent Day-2 subspace occupied a larger representational volume, corresponding to higher memory performance compared to Day-1 clips (70.9 vs. 59.1%). These results suggest that episodic time is encoded as a dynamically maintained neural manifold whose geometric integrity supports accurate memory decisions.