Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms

Reading & Writing Visual Cortex: Sparse Autoencoders as a Semantic Interface to the Brain

Mario Serrafero1, Marlin Lee1, Tiasha Saha Roy1, Daniel Boley1, Thomas Naselaris1; 1University of Minnesota

Presenter: Mario Serrafero

Deep vision models predict cortical responses to natural scenes with striking accuracy, yet their dense feature spaces remain opaque. Sparse autoencoders (SAEs) can factorize these features into interpretable concepts that support causal intervention: they can be modulated at inference time to *steer* the model and alter its outputs. Are these latent concept dimensions driving the alignment with the visual cortex? Do they map to semantically-relevant brain regions, and if so, can we leverage them to edit the representations encoded in brain activity? We investigate by composing an SAE dictionary learned from DINOv2 with ridge encoding weights fit to 7T fMRI (Natural Scenes Dataset), yielding a concept-by-vertex alignment matrix Φ mapping each of 10,000 SAE concepts to its spatial recruitment profile across the cortex. This single object (Φ) supports three applications: (1) *interpretation*–cortical profiles reveal structured, retinotopically organized concept recruitment, with leverage scores ranking concepts by cortical impact; (2) *stimulus retrieval*–correlating measured fMRI with concept profiles retrieves semantically appropriate natural images (brain-side most-eliciting inputs); (3) *causal steering*–modulating concept activations via Φ and decoding through a diffusion model produces targeted semantic edits, demonstrating fine-grained compositional control of neural features. Our results suggest that AI concept directions are causally grounded in cortex representations, bridging mechanistic interpretability and neuroscience.

Topic Area: Computational Models of Vision & Visual Cortex