Keynotes | K&Ts | GACs | Talks | Posters | Search

Contributed Talk Session: Wednesday, August 5, 2:00 – 3:00 pm, Skirball Theater
Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Using Motion-guided Foveation to Improve Physical Prediction in Latent Representation Video Models

Xiangzhou Sun1, Nancy Kanwisher1, RT Pramod1; 1Massachusetts Institute of Technology

Presenter: Xiangzhou Sun

Recent research on video models used latent representations to predict physical interactions in large video datasets. Inspired by the biological visual system, we investigate whether incorporating motion-guided foveation improves physical prediction in video models. Unlike standard approaches that process full-resolution frames uniformly, human vision selectively samples high-acuity information around fixation while compressing peripheral detail, dynamically shifting overt attention through saccades and smooth pursuits. To test whether these constraints aid intuitive physical reasoning, we construct modified versions of the Physion dataset with fixed-center and motion-guided (optical flow) foveation. We evaluate our datasets on R3M+CTRNN and V-JEPA representations using linear and attentive probes. We find that foveation consistently improves representation quality, with fixed-center foveation yielding substantial gains in per-scenario performance (+9.8% AUROC), while optical-flow-driven foveation provides the strongest improvements in holdout generalization (+4.3% AUROC). Additionally, fixed-center and motion-guided foveation more closely resemble human behavior compared to the baseline. These results indicate that reducing peripheral information and guiding foveation toward regions with task-relevant motion contrast can improve physical reasoning, supporting the hypothesis that human-like visual constraints provide computational advantages for dynamic physical predictions.

Topic Area: Computational Models of Vision & Visual Cortex