Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster D in Poster Session D: Wednesday, August 5, 2:00 – 3:45 pm, Kimmel Center, Shorin & Rosenthal Rooms

A Unified Speech-to-Text-to-Speech Encoding Model for Natural Conversations

Haocheng Wang1, Itamar Jalon1, Samuel Nastase2, Ariel Goldstein3, Zaid Zada1, Bobbi Aubrey1, Orrin Devinsky4, Adeen Flinker5, Uri Hasson6; 1Princeton University, 2University of Southern California, 3Hebrew University of Jerusalem, 4NYU Langone, 5New York University, 6Weizmann Institute of Science

Presenter: Haocheng Wang

Humans effortlessly alternate between speaking and listening during everyday conversations. What are the neural mechanisms that underlie these complex behaviors? Here, we investigated the neural basis of speech comprehension and production, leveraging around 100 hours of spontaneous, naturalistic conversations and concurrent electrocorticography (ECoG) recordings. We employed SeamlessM4T, a multimodal large language model (LLM) supporting both speech and text input and output. We extracted (1) speech perception embeddings from the speech encoder, (2) language embeddings from the text decoder, and (3) speech articulation embeddings from the text-to-speech-unit decoder. We then constructed joint electrode-wise encoding models using banded ridge regression. During comprehension, auditory areas were best predicted by speech perception embeddings, while semantic regions were best predicted by language embeddings. During production, the pre- and postcentral gyrus were best predicted by speech articulation and perception embeddings. Notably, the temporal pole was best predicted by speech articulation embeddings. Furthermore, our encoding model enabled us to track the temporal dynamics of speech comprehension and production. Overall, our findings highlight the power of LLMs to capture the neural patterns unfolding in natural conversations: from sensory perception to comprehension, and from planning to motor execution.

Topic Area: Auditory, Speech & Language Processing