Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

Foveated multiview models of human 3D shape perception

Stephanie Fu1, Trevor Darrell2, tyler bonnen1; 1University of California, Berkeley, 2Electrical Engineering & Computer Science Department

Presenter: Stephanie Fu

Two recent developments in computer vision have brought us closer to human-like perception. First, multiview models trained on visual-spatial data have closed a longstanding gap between human and machine 3D perception; remarkably, these models demonstrate an emergent alignment to human error patterns and reaction times. Second, autoregressive gazing models trained for general-purpose reconstruction objectives exhibit an emergent alignment with human fixation patterns - despite no exposure to eye-tracking data. Here we ask whether these developments can be integrated into a foveated multiview transformer. We construct a model without any re-training, determine its zero-shot performance on 3D vision benchmarks, and evaluate its alignment to human behavior. The model retains meaningful task performance across all benchmarks even with sparse, low-resolution inputs, and its gaze patterns correlate with human gaze despite no training on eye-tracking data. These zero-shot results establish foveated multiview models as a promising direction for vision systems that are performant on 3D tasks and grounded in the mechanisms of human visual perception.

Topic Area: Computational Models of Vision & Visual Cortex