Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms
Model manifold analysis suggests the human ventral visual system is less like an optimal classifier and more like a feature bank
Colin Conwell1, Michael Bonner2; 1After Thought, 2Johns Hopkins University
Presenter: Colin Conwell
What do deep neural network (DNN) models actually tell us about the computational principles of visual information-processing in the biological brain? A now common finding in visual neuroscience is that many different kinds of DNNs -- with different architectures, tasks, and training diets -- are all comparably performant predictors of image-evoked brain activity in the ventral visual cortex. This relative parity of diverse models may at first seem to undermine the common intuition that these models can be used to infer the computational principles that govern the visual brain. In this work, we show to the contrary that comparable brain-predictivity does not preclude the differentiation of these same models in terms of the underlying manifold geometries that define them. To do this, we assess 12 manifold geometry metrics computed on a diverse set of 117 DNN models, curated to include multiple tasks, architectures, and input diets. We then use these metrics to predict how well each model aligns with occipitotemporal cortex (OTC) activity from the human fMRI Natural Scenes Dataset. We find that *manifold signal-to-noise ratio* (a metric previously associated with few-shot learning) is a robust predictor of downstream brain-alignment and supersedes both other manifold geometry metrics (i.e. *manifold capacity*) and downstream task-performance (e.g. top-k recognition accuracy) across multiple image sets (e.g. ImageNet21K versus Places365) and controlled model comparisons (e.g. assessments across ImageNet-1K trained architectural variants only). These results add to a growing body of evidence that the ventral visual stream serves as a basis set (or feature vocabulary) for object recognition rather than as the actual locus of recognition *per se*.
Topic Area: Computational Models of Vision & Visual Cortex