Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms

Process-Level Evaluation of Social Inference in LLMs

Elif Akata1, Dylan Cope2, Eric Schulz1, Jakob Nicolaus Foerster2; 1Helmholtz Zentrum München, 2University of Oxford

Presenter: Elif Akata

Humans do not wait until an action sequence ends to infer another agent's goal. They update beliefs online as new evidence arrives. This process can be formalised as Bayesian inverse planning, where an observer maintains a posterior over possible goals and updates it according to each action's diagnostic value. Typical evaluation of Large Language Models (LLMs) on social reasoning through final judgments after a full observed sequence leaves open whether they actually track social evidence over time. We introduce an evaluation framework in which an agent produces actions under a known rational policy in a controlled set of hidden-goal environments. We compute trajectories of ideal-observer posteriors over goals at every step, and evaluate LLMs both on final accuracy and process-level properties of belief updating. Our results show that success on goal inference tasks need not imply human-like tracking of social evidence, and process-level evaluation reveals differences that final-answer metrics obscure.

Topic Area: Methods, Tools, Theory & Neural Coding