Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster B in Poster Session B: Tuesday, August 4, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

CogMaze: A Benchmark of Egocentric Spatial Understanding in Humans and Vision-Language Models

Harvey Donnelly1, Frank Keller1, Benjamin Peters1; 1University of Edinburgh

Presenter: Harvey Donnelly

Humans are adept at spatial understanding: integrating egocentric observations to support navigation and reasoning. Evidence suggests this relies on an allocentric representation of the environment (cognitive maps). While vision-language models (VLMs) perform well on many multimodal tasks, they remain weak at spatial understanding, suggesting they don't form cognitive map-like representations. We introduce CogMaze, a benchmark of spatial understanding in humans and VLMs. It includes a 3D simulated maze with visual landmarks, a ground-truth schema for probing representations, and tasks for humans and VLMs. We evaluate with a VLM in environments of varying complexity.

Topic Area: Methods, Tools, Theory & Neural Coding