Keynotes | K&Ts | GACs | Talks | Posters | Search

Vision

Contributed Talk Session: Wednesday, August 5, 2:00 – 3:00 pm, Skirball Theater

Using Motion-guided Foveation to Improve Physical Prediction in Latent Representation Video Models

Talk 1, 2:00 pm

Xiangzhou Sun1, Nancy Kanwisher1, RT Pramod1; 1Massachusetts Institute of Technology

Presenter: Xiangzhou Sun

Recent research on video models used latent representations to predict physical interactions in large video datasets. Inspired by the biological visual system, we investigate whether incorporating motion-guided foveation improves physical prediction in video models. Unlike standard approaches that process full-resolution frames uniformly, human vision selectively samples high-acuity information around fixation while compressing peripheral detail, dynamically shifting overt attention through saccades and smooth pursuits. To test whether these constraints aid intuitive physical reasoning, we construct modified versions of the Physion dataset with fixed-center and motion-guided (optical flow) foveation. We evaluate our datasets on R3M+CTRNN and V-JEPA representations using linear and attentive probes. We find that foveation consistently improves representation quality, with fixed-center foveation yielding substantial gains in per-scenario performance (+9.8% AUROC), while optical-flow-driven foveation provides the strongest improvements in holdout generalization (+4.3% AUROC). Additionally, fixed-center and motion-guided foveation more closely resemble human behavior compared to the baseline. These results indicate that reducing peripheral information and guiding foveation toward regions with task-relevant motion contrast can improve physical reasoning, supporting the hypothesis that human-like visual constraints provide computational advantages for dynamic physical predictions.

Recycling and Co-option Are Dissociable Mechanisms of Cortical Reorganization

Talk 2, 2:10 pm

Lauren S Aulet1; 1University of Massachusetts at Amherst

Presenter: Lauren S Aulet

When children learn to read, a region of ventral temporal cortex becomes selective for written words, yet this territory evolved long before writing existed. The dominant account of this phenomenon, cortical recycling, proposes that literacy reshapes prior visual representations to serve reading. However, recent work reveals conflicting evidence about which representations are recycled and whether modification is even necessary. We propose that these findings reflect two dissociable mechanisms that have been conflated under a single label. We trained ResNeXt-50 deep neural networks on infant egocentric video (SAYCam) using DINO self-supervised learning and then fine-tuned them on letter classification, tracking channel-level category selectivity. Limb-selective channels show high enrichment among newly letter-selective channels and high feature transformation: they are recycled. Face-selective channels show equally high enrichment but minimal feature transformation: they are co-opted, with a new downstream readout exploiting their existing features. Both mechanisms operate simultaneously on different neural populations during the same learning event, with the determining factor being whether a channel's existing features are already informative for the new task. We propose recycling and co-option as a more precise vocabulary for understanding how the brain repurposes existing circuits for culturally novel skills.

Lateral Recurrence as a Domain-Selective Mechanism

Talk 3, 2:20 pm

Amirhossein Farzmahdi1, Hossein Adeli1, Wang Boran1, Chase King1, Nikolaus Kriegeskorte1; 1Columbia University

Presenter: Amirhossein Farzmahdi

Recurrent connections pervade the primate ventral stream, yet their computational role remains debated. One view holds that recurrence supports iterative inference and can stabilize or refine representations across inputs (Rao & Ballard, 1999; Lee & Mumford, 2003); another suggests that experience shapes recurrence to selectively refine behaviorally relevant information. Face-selective regions exhibit refined identity representations across hierarchical stages (Freiwald & Tsao, 2010; Chang & Tsao, 2017), and recurrent architectures better capture primate visual dynamics (Kar et al., 2019; Kietzmann et al., 2019). However, direct evidence for domain-selective recurrence remains limited because anatomy, experience, and task demands covary in cortex. We disentangle these factors using lateral-recurrent convolutional networks with identical architectures trained on different objectives, face identification versus object categorization. We track representational changes across recurrent time through decoding, population geometry, and single-unit analyses. We find a double dissociation: recurrence strengthens identity information in the trained domain while leaving the untrained domain largely unchanged. This selectivity emerges in deeper layers, scales with recurrence strength, and reflects the alignment of category preference to identity sensitivity at the single-unit level. These results argue that lateral recurrence is a learned, task-specific refinement mechanism and generate testable predictions for time-resolved recordings in domain-selective cortex.

Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex

Talk 4, 2:30 pm

Abdulkadir Gokce1, Yingtian Tang1, Martin Schrimpf1; 1EPFL - EPF Lausanne

Presenter: Abdulkadir Gokce

Task-optimized deep neural networks are the current leading models of visual cortex, but how to build more predictive models remains unclear. We here ask to what extent model-brain alignment is driven by i) scaling generic visual pretraining, ii) fine-tuning with neural supervision, and iii) improving the final mapping from model features to neural responses. We analyze these three levers in a unified pipeline across 600+ vision models trained under controlled conditions, and eight public datasets spanning macaque electrophysiology and human fMRI, EEG, and MEG. Across modalities, scaling pretraining compute and data reliably improve alignment, but gains saturate. In contrast, hybrid task+neural fine-tuning yields consistent improvements that generalize across modalities. The strongest within-dataset gains arise from the mapping stage: increasing the number of paired stimulus--response samples to fit the readout yields robust, near log-linear improvements. Finally, we introduce a subject-shared attention-based readout that matches or exceeds standard per-subject linear probes while using an order of magnitude fewer parameters. Together, these results suggest that progress in modeling the brain will likely come from targeted neural supervision, richer multi-subject mappings, and larger neural datasets.

Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field

Talk 5, 2:40 pm

Dylan Matthew Diaz1, Margaret Marie Henderson2; 1Purdue University, 2Carnegie Mellon University

Presenter: Dylan Matthew Diaz

Within topographic eccentricity maps in the primate visual system, center-preferring populations have higher spatial resolution and overlap face- and word-selective regions while periphery-preferring populations have lower spatial resolution and overlap scene-selective regions. Prior behavioral and neuroimaging evidence suggests that this "eccentricity bias" may reflect the relevance of visual field portions for different tasks: the central visual field may be more informative for fine-grained tasks like face recognition and reading, while the periphery may be more informative for large-scale scene understanding tasks. To examine whether such eccentricity-dependent coding can emerge from natural experience, we leveraged egocentric video and eye-tracking data from the Visual Experience Dataset (VEDB). We trained ResNet-18 models using contrastive learning (SimCLR) on video frames modified to isolate information available at different eccentricities (gaze-contingent fovea-only crops, periphery-only crops, and periphery-only crops with a NeuroFovea transform applied). We then evaluated downstream task performance and model alignment with human fMRI data (Natural Scenes Dataset; encoding model framework). When examining in-domain classification of VEDB frame categories, we observed systematic variability in the performance of fovea-only and periphery-only models across categories, suggesting differential informativeness of visual field eccentricities across tasks. On downstream classification without fine-tuning, VEDB-pretrained models generalized more strongly to scene recognition (Places365) than to face recognition (VGGFace2), with fovea-only models showing an advantage on both tasks. Across visual cortex, VEDB-pretrained models achieved similar neural predictivity to models trained on mid-sized non-egocentric datasets (ImageNet-100), suggesting experience-sampled egocentric data, despite its low diversity and constrained semantic content, supports emergence of cortically-aligned representations. In scene-selective sub-regions (PPA, RSC), periphery-only models held a small but consistent advantage in explained variance over fovea-only models, suggesting scene-selective cortex may be adapted to peripheral-field statistics. Together, these results suggest that naturalistic egocentric experience provides an organizing constraint on perception, leading to adaptive, task-aligned information processing.

Concept Manifold Geometry Explains Asymmetry in Model–Brain Bidirectional Predictivity

Talk 6, 2:50 pm

Hyewon Willow Han1, Qingqing Yang2, Yalda Mohsenzadeh1; 1University of Western Ontario, 2Ohio State University, Columbus

Presenter: Hyewon Willow Han

Strong prediction of brain responses by model features does not guarantee that models capture the true structure of biological visual representations. Moreover, model–brain alignment is not fully characterized by one-directional comparisons; asymmetry between bidirectional mappings may provide a more informative measure. Using activations of deep neural networks and human fMRI data, we systematically examined the bidirectional mappings between the models and human brains, and found forward predictivity consistently exceeding reverse predictivity. Furthermore, we tested whether geometric properties of model concept manifolds account for the variation in forward, reverse predictivity, and their asymmetry. Critically, we found that manifold geometric properties explained substantially more variance in reverse predictivity and asymmetry than in forward predictivity, with Effective Dimensionality, Signal, and Correlation emerging as key predictors. These findings suggest that bidirectional mapping provides a more complete and diagnostic measure of alignment, and manifold geometry accounts for the extent to which the representational properties of models and brains are recoverable from each other. Together, this framework offers a principled path toward models that are more faithfully aligned with the human visual system.