Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms
Predicting Detail from Blur Enables Learning of Robust Visual Representations
Akihito Maruya1, Hossein Adeli1, Tian Zeng1, Nikolaus Kriegeskorte1, Ning Qian1; 1Columbia University
Presenter: Akihito Maruya
The human retina is highly non-uniform, with spatial resolution highest at the fovea and decreasing sharply with eccentricity, despite the subjective impression of a uniformly detailed visual world. We propose that the brain predicts fine peripheral structure from degraded input across saccades, and that this predictive–corrective cycle drives visual learning. Motivated by this, we replace blank tokens in masked autoencoder (MAE) pre-training with blurred tokens that preserve coarse spatial structure, mimicking peripheral sampling. We ask: (Q1) Can a Vision Transformer (ViT) reconstruct full-resolution detail from blurred, incomplete input? (Q2) Can predicting from bandwidth-limited input yield transferable representations that improve classification over blank masking? We pre-trained ViT with blank vs. blur masking across mask ratios (0.65–0.95) on CIFAR-10 and ImageNet, then fine-tuned for classification. Both questions were answered affirmatively. Blur-masked models reconstructed coherent images even at high mask ratios where blank-masked models failed (Q1), and consistently outperformed blank masking in classification — especially under high mask ratios and limited training (Q2). Spatial frequency (SF) analysis revealed that under challenging conditions — high mask ratios and limited training — blank-trained models shift toward high-SF preference, making them susceptible to high-SF representations that are contributing to lower classification accuracy. Blur-trained models, by contrast, maintain low-to-mid SF preference and remain noise resistant across these conditions. These findings draw a natural parallel to biological vision: humans operate under limited foveal access and finite visual experience, and blur masking — like peripheral vision — provides reliable low-frequency context that scaffolds robust learning even under impoverished conditions.
Topic Area: Computational Models of Vision & Visual Cortex