

Gedas Bertasius
Assistant Professor
I am an Assistant Professor in the Computer Science department at the University of North Carolina, Chapel Hill. Before joining UNC, I was a postdoctoral researcher at Facebook AI Research (FAIR) working with Lorenzo Torresani. I received my Ph.D. at the University of Pennsylvania, where I was advised by Jianbo Shi, and my undergraduate degree at Dartmouth College.
Research
I lead the Multimodal Video Perception (MVP) group at UNC. The long-term goal of our research is to develop machines that understand human behavior from video. We pursue this goal across three levels: (1) building scalable representations of long, multimodal video; (2) developing methods that infer the intentions, capabilities, and decisions behind observed behavior; and (3) designing systems that translate these inferences into coaching, task assistance, and embodied action. Representative projects include TimeSformer, Video ReCap, LLoVi, VideoTree, ExAct, SVI-Bench, and WatchAct.
Video Recognition

Designing the core architectures for video recognition, from short actions to long-form activities (e.g., TimeSformer, ViS4mer).
Multimodal AI
Building models that jointly understand video, audio, and language, from hour-long captioning to long-range question answering (e.g., Video ReCap, LLoVi, BIMBA).

Perceptual Assistants & Coaches
Strategic Video Intelligence

Generative Video Modeling


Teaching robots to act by watching people, from behavior-grounded manipulation to robust long-horizon execution (e.g., WatchAct, BOSS, ReBot, and ARCADE)
Selected Projects
SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence
Yulu Pan, Han Yi, Seongsu Ha, Md Mohaiminul Islam, Benjamin Zhang, Lorenzo Torresani, Gedas Bertasius
ECCV 2026
[arxiv] [video] [project page] [extended paper] [code] [data] [bibtex]
WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation
Baiqi Li, Ce Zhang, Yu Fang, Yue Yang, Shangzhe Li, Mingyu Ding, Gedas Bertasius
arXiv 2026
[arxiv] [project page] [code] [data] [bibtex]
SiLVR: A Simple Language-based Video Reasoning Framework
Ce Zhang, Yan-Bo Lin, Ziyang Wang, Mohit Bansal, Gedas Bertasius
TMLR 2026 (1st Place, CVPR Multi-Discipline Lecture Understanding Challenge)
[arxiv] [project page] [code] [bibtex]
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives
Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Gedas Bertasius, ... , Michael Wray
CVPR 2024
Video ReCap: Recursive Captioning of Hour-Long Videos
Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, Gedas Bertasius
CVPR 2024 (Egocentric Vision Distinguished Paper Award)
[arxiv] [project website] [code] [dataset] [bibtex]
Is Space-Time Attention All You Need for Video Understanding?
Gedas Bertasius, Heng Wang, Lorenzo Torresani
ICML 2021 (Top-5 Most Cited ICML 2021 Paper)
[arxiv] [code] [talk] [slides] [blog] [VentureBeat] [SiliconAngle] [bibtex]
Sponsors
We are grateful for the following agencies for supporting our research.
















