Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

The Transformer May Be the Universal Cortical Microcircuit

Yash Shah1, Daniel LK Yamins1; 1Stanford University

Presenter: Yash Shah

The primate ventral visual system is constrained by the dual evolutionary pressures of (i) enabling strong visual behavior, as observed via high-accuracy readout of object and action categories, and (ii) being wiring efficient. These pressures have shaped a cortical architecture with putative hierarchical linear-nonlinear processing layers, systematic growth of receptive fields across areas, and functional organization of cortical cells. A key challenge has been to identify the cortical microcircuit that satisfies the dual constraints. The convolution---as implemented by convolutional neural networks (CNNs)---has served as the de facto model of ventral stream microcircuitry. However, previous works have shown that they display a small drop in task performance whilst trying to achieve efficient wiring. Here, we propose that the transformer implements a more complete model of the cortical microcircuit. Although transformers are commonly assumed to rely on global, long-range intra-layer interactions, we show that vision transformers in fact develop a hierarchy of receptive field sizes with depth. The transformer microcircuit additionally achieves efficient feedforward and intra-layer wiring without sacrificing task performance. Critically, what drives this alignment is not the existence of global dependencies but a multiplicative operation, as operationalized via the attention module, in addition to additive operations also present in the CNN.

Topic Area: Computational Models of Vision & Visual Cortex