Keynotes | K&Ts | GACs | Talks | Posters | Search

Poster F in Poster Session F: Thursday, August 6, 1:45 – 3:30 pm, Kimmel Center, Shorin & Rosenthal Rooms

VAST: Visual Abstractions as Streams of Thought

Khai Loong Aw1, Klemen Kotar1, Baihan Zhang1, Yifan Liu2, Lilian Naing Chen1, Rahul Mysore Venkatesh1, Atlas Kazemian1, Wanhee Lee1, Daniel LK Yamins1; 1Stanford University, 2Tsinghua University

Presenter: Khai Loong Aw

Visual understanding is as fundamental as language, anchoring how we reason about the physical world. Yet modern vision-language models (VLMs) reason primarily by producing language tokens and rely on pretrained language backbones consuming most parameters and compute. We introduce Visual Abstractions as Streams of Thought (VAST), a cognitively-motivated VLM that reasons through visual streams of thought: lower-dimensional abstractions (camera pose, depth, optical flow, point tracks, and text) disentangled from appearance and interpretable when predictions fail. VAST offers two capabilities beyond standard VLMs. First, richly controllable world modeling: conditioning on any combination of modalities (e.g., optical flow, depth, text, point tracks) with patch-level independence for targeted object manipulations, enabling fine-grained counterfactual reasoning. Second, a single model unifies diverse visual-linguistic tasks, e.g., novel view synthesis, captioning, action-conditioned generation, by expressing each as a stream of thought. Together, VAST offers a path toward reasoning that is grounded in the physical world.

Topic Area: Computational Models of Vision & Visual Cortex