Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster E in Poster Session E: Thursday, August 6, 10:30 am – 12:15 pm, Kimmel Center, Shorin & Rosenthal Rooms
Transitivity can redeem the reversal curse in attention models
Yedi Zhang1, Andrew Kyle Lampinen2, James McClelland3; 1University College London, 2Google DeepMind, 3DeepMind, London, UK
Presenter: Yedi Zhang
Despite their remarkable abilities, attention-based language models show a basic generalization failure known as the reversal curse: they fail to infer "A -> B" after training on "B <- A". We study this phenomenon in a minimal two-layer linear attention model trained on a tournament task to predict binary labels for entity pairs. The model exhibits the reversal curse when the tournament is intransitive, but overcomes it when the tournament is transitive (i.e., if A -> B and B -> C, then A -> C). This behavioral difference can be explained by a representational difference. We derive a reduction that decomposes the model's computation into two distinct components: an additive term that ranks the entities, and a conjunctive term that memorizes specific pairs. The model relies on additive representations that support relational generalization in the transitive case, while it relies on conjunctive representations that do not support generalization in the intransitive case. These results show how transitivity can redeem the reversal curse while clarifying why many intransitive yet reversible relations remain cursed.
Topic Area: Methods, Tools, Theory & Neural Coding