Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms

Remembered dialogue can shift alignment-relevant behavior in LLMs, with conditional effects of source credibility

Chad DeChant1, Marcelo G Mattar1; 1New York University

Presenter: Chad DeChant

Large language models are trained to satisfy behavioral and safety constraints, but this alignment is not always robust. Recent work has shown that fine-tuning on incorrect question-answer pairs produces broad shifts toward dangerous or inaccurate outputs, a phenomenon called emergent misalignment. Whether analogous effects can arise at inference time, without weight updates, and what factors modulate such susceptibility, remain unclear. Here we show that presenting as few as three examples of flawed dialogue exchanges framed as recalled memories of a user's past conversations can induce misalignment in several widely used models. Critically, when the model is given weaker instructions to merely ``take into account'' these memories, the social role attributed to the user's conversation partner significantly modulates the effect, consistent with sensitivity to source credibility cues. These findings suggest that memory recall in realistic deployment scenarios poses alignment risks that depend not only on content but on social context.

Topic Area: Auditory, Speech & Language Processing