Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

Learning to see via predictive exploration of a 3D environment enhances visual representation of scene geometry

Rupert Tawiah-Quashie1, Richard Hakim1, George A. Alvarez1, Talia Konkle1; 1Harvard University

Presenter: Rupert Tawiah-Quashie

In seminal work, Held & Hein (1963) showed that active visual exploration, versus passive viewing, is essential for normal visual development, implicating vision-action integration as a key driver of representation learning. Here, we ask whether predictive navigation training shapes visual feature representations in a world model (tiny-JEPA: ConvNet backbone + LSTM), trained end-to-end to predict its own latent visual responses given its current latent and actions (dx, dy, dθ). After training, we evaluated CNN and RNN latent representations using linear probes to decode viewpoint from frozen weights. Following environment warmup, agents performed a structured rotation sequence in different environments. In environments where individual frames were geometrically ambiguous (16 evenly-spaced steps, preserving 4-fold room symmetry), only the LSTM's recurrent state supported viewpoint decoding, consistent with temporal integration of heading. In environments with subtle but distinct per-view scene geometry (21 steps, breaking room symmetry), the navigation-trained CNN backbone also supported decoding, while a similar AlexNet architecture trained on ImageNet performed near chance. These results suggest that predictive navigation training drives the emergence of geometry-sensitive visual features absent in standard object-recognition models, and that recurrent world-model states encode sequential heading information that the backbone alone cannot provide

Topic Area: Computational Models of Vision & Visual Cortex