Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster C in Poster Session C: Wednesday, August 5, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Modeling Human Reasoning Errors on ARC with Natural Language Programs

Jinran Jin1, Solim LeGris1, Todd M. Gureckis1, Brenden Lake2; 1New York University, 2Princeton University

Presenter: Jinran Jin

The Abstraction and Reasoning Corpus (ARC) is a challenging benchmark for evaluating AI systems on inferring rule-like concepts from just a few examples. Humans solve ARC tasks with little to no specialized training, while AI systems have only begun to match humans despite extensive compute and data scaling. We study what makes humans such sample-efficient reasoners by modeling their mistakes, as cataloged in the H-ARC dataset. We hypothesize that people solve ARC problems with internal representations that are best captured with natural language "programs". To examine this hypothesis, we first generate natural language programs using an LLM generator, and validate them with an LLM interpreter that executes them. We then systematically perturb these validated programs. Overall we reproduced errors covering 13.4% of human errors in the H-ARC dataset. Together, our findings suggest that natural language programs provide a viable computational framework for modeling human abstract reasoning, and that human errors are best explained as incomplete or approximate rule representations rather than random failure.

Topic Area: Memory, Learning & Knowledge Structures