Keynotes | K&Ts | GACs | Talks | Posters | Search
Contributed Talk Session: Wednesday, August 5, 2:00 – 3:00 pm, Skirball Theater
Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms
Multimodal Scaling Laws for Task & Data-Optimized Models of Visual Cortex
Abdulkadir Gokce1, Yingtian Tang1, Martin Schrimpf1; 1EPFL - EPF Lausanne
Presenter: Abdulkadir Gokce
Task-optimized deep neural networks are the current leading models of visual cortex, but how to build more predictive models remains unclear. We here ask to what extent model-brain alignment is driven by i) scaling generic visual pretraining, ii) fine-tuning with neural supervision, and iii) improving the final mapping from model features to neural responses. We analyze these three levers in a unified pipeline across 600+ vision models trained under controlled conditions, and eight public datasets spanning macaque electrophysiology and human fMRI, EEG, and MEG. Across modalities, scaling pretraining compute and data reliably improve alignment, but gains saturate. In contrast, hybrid task+neural fine-tuning yields consistent improvements that generalize across modalities. The strongest within-dataset gains arise from the mapping stage: increasing the number of paired stimulus--response samples to fit the readout yields robust, near log-linear improvements. Finally, we introduce a subject-shared attention-based readout that matches or exceeds standard per-subject linear probes while using an order of magnitude fewer parameters. Together, these results suggest that progress in modeling the brain will likely come from targeted neural supervision, richer multi-subject mappings, and larger neural datasets.
Topic Area: Computational Models of Vision & Visual Cortex