Keynotes | K&Ts | GACs | Talks | Posters | Search

Contributed Talk Session: Wednesday, August 5, 2:00 – 3:00 pm, Skirball Theater
Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms

Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field

Dylan Matthew Diaz1, Margaret Marie Henderson2; 1Purdue University, 2Carnegie Mellon University

Presenter: Dylan Matthew Diaz

Within topographic eccentricity maps in the primate visual system, center-preferring populations have higher spatial resolution and overlap face- and word-selective regions while periphery-preferring populations have lower spatial resolution and overlap scene-selective regions. Prior behavioral and neuroimaging evidence suggests that this "eccentricity bias" may reflect the relevance of visual field portions for different tasks: the central visual field may be more informative for fine-grained tasks like face recognition and reading, while the periphery may be more informative for large-scale scene understanding tasks. To examine whether such eccentricity-dependent coding can emerge from natural experience, we leveraged egocentric video and eye-tracking data from the Visual Experience Dataset (VEDB). We trained ResNet-18 models using contrastive learning (SimCLR) on video frames modified to isolate information available at different eccentricities (gaze-contingent fovea-only crops, periphery-only crops, and periphery-only crops with a NeuroFovea transform applied). We then evaluated downstream task performance and model alignment with human fMRI data (Natural Scenes Dataset; encoding model framework). When examining in-domain classification of VEDB frame categories, we observed systematic variability in the performance of fovea-only and periphery-only models across categories, suggesting differential informativeness of visual field eccentricities across tasks. On downstream classification without fine-tuning, VEDB-pretrained models generalized more strongly to scene recognition (Places365) than to face recognition (VGGFace2), with fovea-only models showing an advantage on both tasks. Across visual cortex, VEDB-pretrained models achieved similar neural predictivity to models trained on mid-sized non-egocentric datasets (ImageNet-100), suggesting experience-sampled egocentric data, despite its low diversity and constrained semantic content, supports emergence of cortically-aligned representations. In scene-selective sub-regions (PPA, RSC), periphery-only models held a small but consistent advantage in explained variance over fovea-only models, suggesting scene-selective cortex may be adapted to peripheral-field statistics. Together, these results suggest that naturalistic egocentric experience provides an organizing constraint on perception, leading to adaptive, task-aligned information processing.

Topic Area: Computational Models of Vision & Visual Cortex