Keynotes | K&Ts | GACs | Talks | Posters | Search
Contributed Talk Session: Thursday, August 6, 1:45 – 2:45 pm, Skirball Theater
Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms
Vision Models Capture Complementary Representations in Children That Converge in Adults
Domenic Bersch1, Timothy Schaumlöffel1, Michela Proietti1, Antonia Franaszek-Traczewska2, Hannah Elisabeth Zwad2, Siying Xie2, Marlena Baldauf3, Teresa Sylvester2, Christina Maria Schätz4, Moritz Köster3, Stefanie Höhl4, Bert Turtleton2, Radoslaw Martin Cichy2, Gemma Roig1; 1Johann Wolfgang Goethe Universität Frankfurt am Main, 2Freie Universität Berlin, 3University of Regensburg, 4University of Vienna
Presenter: Domenic Bersch
How the relationship between image-text aligned and vision-only models changes across development remains unknown. To address this, we applied time-resolved representational similarity analysis to EEG data from 8-year-olds, 12-year-olds, and adults. Commonality analysis showed complementary structure for image-text aligned versus vision-only models at 8 years, near-zero shared variance at 12 years, and greater overlap in adults, while both retained unique variance at all ages. This pattern generalized across model pairings, whereas vision-only model pairs showed overlapping variance at every age, suggesting a double dissociation. Within-category structure appeared complementary even in adults, suggesting that image-text aligned training may capture fine-grained structure not recovered by vision-only training. Together, these findings suggest that image-text aligned and vision-only models capture distinct aspects of visual representations that become more aligned in adulthood, and that vision-only models alone may miss structure relevant for developmental model-brain alignment.
Topic Area: Development, Individual Differences & Clinical Populations