Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Simple 3D Pose Features Support Human and Machine Social Scene Understanding

Wenshuo Qin1, Leyla Isik1; 1Johns Hopkins University

Presenter: Wenshuo Qin

Humans effortlessly recognize social interactions from vision, yet the underlying computations remain poorly understood and challenge even advanced deep neural networks (DNNs). Here, we hypothesized that humans rely on 3D visuospatial pose information to make social judgments, which is largely absent from most DNNs. To test this, we used a novel depth-aware pose estimation pipeline to extract 3D body joints from short video clips and found that they predicted human social judgments better than most embeddings from over 350 vision DNNs. We then reduced these joints to a minimal feature set describing only the 3D position and direction of people and found that this set, but not its 2D counterpart, was necessary and sufficient to match the full body joints' prediction performance. These minimal 3D features also predicted the extent to which DNNs aligned with human social judgments. Together, our findings suggest that human social perception relies on simple, explicit 3D pose information.

Topic Area: Computational Models of Vision & Visual Cortex