Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster C in Poster Session C: Wednesday, August 5, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

The Baby Intuitions Benchmark 2: Evaluating Social Intelligence in Humans and Machines

Shannon Yasuda1, Mark K Ho1, Moira Rose Dillon1; 1New York University

Presenter: Shannon Yasuda

From early in development, the ability to infer the hidden social and object goals driving others’ actions is key to human social learning. Can artificial intelligence infer such goals? To test this question, we present here the Baby Intuitions Benchmark 2 (BIB2) with preliminary data from N = 74 adults, validating expected patterns of human performance. We introduce three computational models to be evaluated on BIB2: (1) a video transformer model shown to succeed on BIB2’s predecessor; (2) a variant of the transformer model that incorporates inductive biases for social learning via pretraining on SAYCam-S, a naturalistic dataset of head-mounted video from a single infant; (3) Google Gemini, a large multimodal language model. By probing humans’ and models’ reasoning about affiliation and agency from minimal cues, BIB2 provides a principled framework for studying the foundations of social intelligence and for advancing human-like AI.

Topic Area: Development, Individual Differences & Clinical Populations