Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms
Temporal Prediction Produces Viewpoint-Consistent, Human-Aligned Vision in DNNs
Peisen Zhou1, Drew Linsley1, Akash Nagaraj1, Alekh Karkada Ashok1, Thomas Serre1; 1Brown University
Presenter: Peisen Zhou
A long tradition in computational neuroscience proposes that the temporal structure of visual experience shapes how biological systems learn to see. Deep neural networks (DNNs), typically trained on static labeled images, lack this inductive pressure. Could the learning regime explain their growing misalignment with human vision? We fine-tuned large pretrained vision transformers on naturalistic videos of real-world objects, systematically comparing self-supervised objectives that range from reconstructing the current frame to interpolating between frames to predicting the next frame. Next-frame prediction in pixel space produced the largest gain, raising alignment with humans to near the human ceiling while maintaining classification accuracy and near-human-level viewpoint stability. These results offer a step toward identifying the learning principles that wire up biological vision.
Topic Area: Computational Models of Vision & Visual Cortex