Keynotes | K&Ts | GACs | Talks | Posters | Search
Poster A in Poster Session A: Tuesday, August 4, 9:30 – 11:15 am, Kimmel Center, Shorin & Rosenthal Rooms
Learning Language by Listening: A Computational Learnability Account
Greta Tuckute1, Klemen Kotar2, Daniel LK Yamins2, Talia Konkle1; 1Harvard University, 2Stanford University
Presenter: Greta Tuckute
Humans learn language from continuous speech, yet most large language models learn from pre-tokenized text and therefore bypass the problem of extracting linguistic structure from raw acoustic input. We present AuriStream Multi-Token Prediction, a fully textless self-supervised model trained to predict future acoustic inputs from continuous speech. By systematically varying the model’s prediction horizon (5–400 ms), we ask what timescale of prediction over continuous speech bootstraps linguistic structure. In these models, lexical abilities (word knowledge) are strongest at intermediate horizons around 200 ms, whereas lower-level acoustic abilities such as pitch are not. These results identify a privileged temporal scale of prediction as a key pressure for learning words from speech and position AuriStream as a perceptually grounded model system for studying the emergence of language from continuous input.
Topic Area: Auditory, Speech & Language Processing