Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster C in Poster Session C: Wednesday, August 5, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Human-level 3D perception emerges from multi-view learning

tyler bonnen1, Jitendra Malik2, Angjoo Kanazawa2; 1University of California, Berkeley, 2Amazon

Presenter: tyler bonnen

Here we introduce the first modeling framework that matches human performance on 3D visual tasks. We leverage a novel class of neural networks trained with a visual-spatial objective; given images from different locations within a scene, these models must learn to predict associated spatial information, such as camera position and visual depth. Notably, these visual-spatial data are analogous to sensory cues readily available to humans. We design a zero-shot evaluation approach to determine the performance of these `multi-view' models on a well established 3D perception benchmark, then compare model and human behavior. Our modeling approach is the first to match human accuracy on 3D shape inferences, even without training or fine-tuning on experimental data or images. Remarkably, independent readouts of model responses predict fine-grained measures of human behavior, including error patterns and reaction times, revealing a natural correspondence between model dynamics and human perception. Our findings indicate that human-level 3D perception can emerge from a simple, scalable learning objective over naturalistic visual-spatial data.

Topic Area: Computational Models of Vision & Visual Cortex