Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms
Blame via Inverse Reinforcement Learning
Eivinas Butkus1, Christian Mott1, Nikolaus Kriegeskorte1, Christopher Baldassano1; 1Columbia University
Presenter: Eivinas Butkus
When someone causes harm, moral judgments depend on inferences about their mental states, including their values. Here we formalize this process using inverse reinforcement learning (IRL). We built a two-agent grid-world simulating people navigating around each other on a street and trained a transformer-based policy conditioned on an "empathy" parameter _ω_ that scales the other agent's reward in the agent's own objective. We then performed IRL to infer _ω_ from observed actions. In a behavioral experiment (_N_ = 21), participants watched 3D animations of agent interactions, inferred the agent's empathy, and then made moral evaluations (e.g., blame) about a later, unrelated harm involving the same agent. The IRL model's trial-level confidence predicted how accurately humans classified the agent's empathy. A mixture of the trial-specific model posterior and the empirical human prior best explained human inferences. Critically, participants evaluated agents with low empathy as significantly worse on multiple moral measures. These results suggest moral evaluations rely, in part, on intuitive value inferences well-captured by IRL.
Topic Area: Decision-Making, Cognitive Control & Event Cognition