
The long-term goal of our research is to develop machines that understand human behavior from video. We build models that perceive long, multimodal videos and reason about why events unfold and what may happen next. Core projects include TimeSformer, Video ReCap, LLoVi, BIMBA, VideoTree.
Building on this foundation, we pursue four directions:
-
PERCEPTUAL ASSISTANTS & COACHES: Guiding people through everyday tasks and coaching skill learning (e.g., VidAssist, Ego-Exo4D, and ExAct).
-
STRATEGIC VIDEO INTELLIGENCE: Reasoning about goals, decisions, and outcomes in dynamic multi-agent settings, with team sports as our testbed (e.g., SVI-Bench, BASKET).
-
ROBOTICS: Teaching robots to act by watching people, from behavior-grounded manipulation to robust long-horizon execution (e.g., WatchAct, BOSS, ReBot, and ARCADE).
-
GENERATIVE VIDEO APPLICATIONS: Generating and editing video and audio together, including video-to-music generation, audio-visual editing, and temporally-consistent video generation (e.g., V2M-Zero, VMAs, AvED, and TeDiO).
Group Photos

































