Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Learning the paths to goals in the absence of reward

Yunchang Zhang1, Yotam Sagiv1, Stefan Oline1, Misael Magos1, Selena Xiang1, Robert Fetcho1, Yousuf El-Jayyousi1, Shai Lipkin1, Nathaniel D. Daw1, Ilana Witten2; 1Princeton University, 2Princeton University; Howard Hughes Medical Institute

Presenter: Yunchang Zhang

Learning the paths to goals is essential for animal survival. While much work has focused on how animals learn based on direct reinforcement, how animals learn to navigate in the absence of reward remains less clear. Here, we developed a novel behavioral task in which mice learn the shortest path to multiple goal locations in a maze during a pre-exposure period without reward. Learning is assessed during a subsequent test period in which mice navigate to rewards efficiently based on their pre-exposure experience. We initially hypothesized that striatal neurons would represent learned spatial maps as a set of value functions over allocentric space for different goals (akin to the successor representation). Using Neuropixels 2.0, we recorded in the nucleus accumbens (NAc), which has been implicated in encoding spatial value, and the dorsomedial striatum (DMS), which has been implicated in goal-directed behavior. We observed distinct neural representations in these two regions: NAc neurons firing ramps as the animal approaches goals (value-like signal), while DMS neurons exhibit sequential activity. Different from our hypothesis, the value representation in NAc is invariant to the current goal identity. A subset of DMS neurons instead appears to encode goal-specific trajectories, but with a particular pattern: they encode the conjunction of proximity to goal and turning angle across trajectories. These data suggest that the brain represents learned spatial policies in a different manner than previously assumed in models of spatial navigation, by decomposing the navigation problem into two components: learning the proximity to each goal and learning what action should be taken at each distance. Such an egocentric representation may be useful for generalizing and reusing learned goal approach policies across many different situations, such as different maze configurations in this task.

Topic Area: Memory, Learning & Knowledge Structures