Keynotes | K&Ts | GACs | Talks | Posters | Search
Learning from depth alone yields strong, shape-based object recognition and alignment with high-level visual cortex
Zejin Lu1, Boyan Rong1, Alexander Kroner2, Daniel Janini1, Radoslaw Martin Cichy1, Tim C Kietzmann2; 1Freie Universität Berlin, 2Universität Osnabrück
Presenter: Zejin Lu
Vision evolved to support action in a structured world, where depth provides critical information about surfaces, boundaries, and spatial layout. Detecting depth structure may therefore be able to guide behaviour independent of recognising fine-grained appearance patterns that appear on top of the depth-defined structure. This raises a key question: can object recognition be supported by depth information alone independent of appearance-based visual properties? To test this idea, we convert natural images into estimated depth maps and train recognition models using depth-only inputs. We observe that models operating over depth maps achieve categorization accuracy comparable to RGB models, exhibit stronger shape bias, and produce representations that align with mid/high-level visual cortex in the Natural Scenes Dataset. These findings show that depth estimates alone provide sufficient information to support object recognition and capture aspects of higher-level visual representations.
Topic Area: Computational Models of Vision & Visual Cortex