Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Othello-GPT Does Not Have a World Model: Lessons for Attributing World Models to Neural Systems

John Morrison1, Mouad Oulouali2, Teo Maayan2; 1Barnard College, 2Columbia University

Presenter: John Morrison

When does a neural system use a world model rather than a collection of heuristics? Othello-GPT is often taken to be a clear and compelling example of a neural system using a world model (see starred references for examples). The evidence falls into three categories: near-perfect prediction of legal moves, near-perfect linear decoding of the board state, and highly successful causal interventions on that board state. We argue that the prediction evidence is explained equally well by heuristics, the decoding evidence relies on the wrong benchmark, and the intervention evidence is inflated by the methodology. As a result, none of this evidence favors one hypothesis over the other. We then introduce a fourth kind of evidence: transfer learning speeds on corrupted versions of Othello. This evidence favors the hypothesis that Othello-GPT is just using a collection of heuristics. We end by proposing a general framework that attributes world models using transfer learning speeds rather than performance, decoding, and interventions.

Topic Area: Methods, Tools, Theory & Neural Coding