Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms

Unable to Forget: Proactive Interference Reveals Working Memory Limits in Large Language Models

Chupei Wang1, Jiaqiu Vince Sun2; 1GiantFish Intelligent Frontier Exploration LLC, 2New York University

Presenter: Jiaqiu Vince Sun

Large language models (LLMs) are increasingly used in settings that require retrieving and updating information from long contexts, yet most evaluations conflate two distinct retrieval challenges: search---locating a target in a vast context---and interference---identifying it when competing, outdated bindings to the same cue are present. Inspired by the proactive interference (PI) paradigm in cognitive science, where earlier associations for a repeated cue disrupt recall of newer ones, we introduce PI-LLM, an evaluation in which interleaved key--value updates are streamed and models must retrieve only the most recent value for each key. Because the target is always the last-presented binding, search difficulty is minimal and interference is isolated as the primary variable. Across 35+ models---from 0.6B open-weight to frontier-scale systems---retrieval accuracy declines approximately log-linearly as same-key updates accumulate, with errors dominated by outdated values. This degradation persists when total input length is held constant and emerges independently across multiple load dimensions, pointing to a working-memory-like bottleneck. Critically, neither explicit 'forget' instructions nor chain-of-thought reasoning alleviates the decline, revealing a dissociation between analytical reasoning and retrieval execution. Unlike the plateau observed in human PI studies, the LLM decline shows no sign of leveling off, suggesting that current models lack the flexible executive control that supports human resilience to interference.

Topic Area: Memory, Learning & Knowledge Structures