top of page
UNC_Zoom_Backgrounds_C.jpg

Gedas Bertasius

Assistant Professor

x-social-media-black-icon.png

I am an Assistant Professor in the Computer Science department at the University of North Carolina, Chapel Hill. Before joining UNC, I was a postdoctoral researcher at Facebook AI Research (FAIR) working with Lorenzo Torresani. I finished my Ph.D. at the University of Pennsylvania, advised by Jianbo Shi, and my undergraduate degree at Dartmouth College.

Research

I lead the Multimodal Video Perception (MVP) group at UNC. The long-term goal of our research is to develop machines that understand human behavior from video. We build models that perceive long, multimodal videos, reason about why events unfold and what may happen next, and use these inferences for coaching, task assistance, and robotic action. Representative projects include TimeSformer, Video ReCap, LLoVi, BIMBA, VideoTree.

Video Recognition

Developing spatiotemporal models for automatic video analysis (e.g., TimeSformer, ViS4mer)

Multimodal AI

Building models that can learn from video, audio, and text (e.g., Video ReCap, LLoVi, BIMBA).

Perceptual AI Coaches

Sports & AI

Developing AI models that can assist people with daily tasks and skill learning (e.g., VidAssist, Ego-Exo4D, and ExAct).

Video for Robotics

Elevating strategic insights using state-of-the-art multimodal video models (e.g., SVI-BenchBASKET).

Generative Video Modeling

Translating visual inputs into effective real-world actions (e.g., WatchAct, BOSS, ReBot, and ARCADE)

Enabling applications such as multimodal video generation and editing (e.g., VMAsAvEDV2M-Zero, and TeDiO

Recent News

Contact

Selected Projects

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

Baiqi Li, Ce Zhang, Yu Fang, Yue Yang, Shangzhe Li, Mingyu Ding, Gedas Bertasius

arXiv 2026

[arxiv] [project page] [code] [data] [bibtex​​​

SiLVR: A Simple Language-based Video Reasoning Framework

Ce Zhang, Yan-Bo Lin, Ziyang Wang, Mohit Bansal, Gedas Bertasius

       TMLR 2026 (1st Place, CVPR Multi-Discipline Lecture Understanding Challenge)

[arxiv] [project page] [code] [bibtex

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Gedas Bertasius, ... , Michael Wray

CVPR 2024

[arxiv] [project website] [blog] [video] [bibtex

Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, Gedas Bertasius

      CVPR 2024 (Egocentric Vision Distinguished Paper Award)

[arxiv] [project website] [code] [dataset[bibtex

Is Space-Time Attention All You Need for Video Understanding?

Gedas Bertasius, Heng Wang, Lorenzo Torresani

      ICML 2021 (Top-5 Most Cited ICML 2021 Paper)

[arxiv] [code] [talk] [slides] [blog] [VentureBeat] [SiliconAngle] [bibtex]

Sponsors

We are grateful for the following agencies for supporting our research.

Contact

Prospective Graduate Students: I am recruiting motivated students in computer vision. Please email me a list of your prior publications and your CV.

Undergraduates at UNC: If you are interested in computer vision, especially its applications to sports, email me your CV and transcript with your GPA.

©2024 by Gedas Bertasius

bottom of page