Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Large-scale vision models show sparks of visual reasoning from natural image training

Brian Kim1, Siyu Zeng1, Drew Linsley1, Alekh Karkada Ashok1, Akash Nagaraj1, Thomas Serre1; 1Brown University

Presenter: Thomas Serre

Large-scale pretraining has become the dominant paradigm in computer vision, but do the representations it produces support visual reasoning—the ability to trace contours, group elements, and resolve spatial relationships? These capabilities have proven difficult to learn directly: neither CNNs nor transformers could solve benchmarks like Pathfinder when trained directly on millions of examples (Kim et al., 2020; Linsley et al., 2018; Tay et al., 2021). Here, we ask whether large-scale pretraining on natural images succeeds where supervised training on these tasks does not. We probe 1,259 pretrained models from the TIMM library on Pathfinder and cluttered ABC (cABC), using linear probes on frozen representations. Most models fail as the required reasoning range increases. However, the largest and most recent ViTs break this trend, showing a strengthening relationship between ImageNet accuracy and visual reasoning performance. These results suggest that large-scale pretraining may give rise to long-range visual reasoning capabilities that explicit training could not—with implications for architecture design and the use of deep networks as models of biological vision.

Topic Area: Computational Models of Vision & Visual Cortex