top of page

Video Recognition

An estimated 3.1 billion people watch Internet video daily, making it one of the largest and richest data sources in existence. Our group develops the core spatiotemporal architectures for analyzing this data, from models of short actions to representations of hour-long recordings.

Related Publications:

Long Movie Clip Classification with State-Space Video Models

Md Mohaiminul Islam, Gedas Bertasius

ECCV 2022

[arxiv] [code] [bibtex]

Long-Short Temporal Contrastive Learning of Video Transformers

Jue Wang, Gedas Bertasius, Du Tran, Lorenzo Torresani

CVPR 2022

[arxiv] [bibtex]

Is Space-Time Attention All You Need for Video Understanding?

Gedas Bertasius, Heng Wang, Lorenzo Torresani

      ICML 2021 (Top-5 Most Cited ICML 2021 Paper)

[arxiv] [code] [talk] [slides] [blog] [VentureBeat] [SiliconAngle] [bibtex]

Multimodal AI

Humans understand the world by combining sight, sound, and language. We build models that do the same, jointly reasoning over video, audio, speech, and text to understand complex real-world content.

Related Publications:

TimeBlind: A Spatio-Temporal Compositionality Benchmark for Video LLMs

Baiqi LiKangyi ZhaoCe ZhangChancharik MitraJean de Dieu Nyandwi

Gedas Bertasius

arXiv 2026

[arxiv] [project page] [code] [dataset] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

Md Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang, Gedas Bertasius, Lorenzo Torresani

      CVPR 2025 (1st Place, CVPR EgoSchema Challenge)

[arxiv] [project page] [code] [model] [demo] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​

Video ReCap: Recursive Captioning of Hour-Long Videos

Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, Gedas Bertasius

      CVPR 2024 (Egocentric Vision Distinguished Paper Award)

[arxiv] [project website] [code] [dataset[bibtex

A Simple LLM Framework for Long-Range Video Question-Answering

Ce Zhang, Taixi Lu, Md Mohaiminul Islam, Ziyang Wang, Shoubin Yu, Mohit Bansal, Gedas Bertasius

EMNLP 2024

[arxiv] [code] [bibtex

Vision Transformers are Parameter-Efficient Audio-Visual Learners

Yan-Bo Lin, Yi-Lin Sung, Jie Lei, Mohit Bansal, Gedas Bertasius

CVPR 2023

[arxiv] [code] [project page] [bibtex

Perceptual Assistants & Coaches

Our group develops perceptual AI agents that help people with daily tasks and skill learning. Our work in this area includes modeling human behavior from first-person video, assisting people with procedural action planning, and understanding and coaching human skills from video.

Related Publications:

ExAct: A Video-Language Benchmark for Expert Action Analysis

Han Yi, Yulu Pan, Feihong He, Xinyu Liu, Benjamin Zhang, Oluwatumininu Oguntola, Gedas Bertasius

NeurIPS Datasets and Benchmarks Track 2025

[arxiv] [project page] [code] [dataset] [leaderboard] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

Propose, Assess, Search: Harnessing LLMs for Goal-Oriented Planning in Instructional Videos

Md Mohaiminul Islam, Tushar Nagarajan, Huiyu Wang, Fu-Jen Chu, Kris Kitani, Gedas Bertasius, Xitong Yang

ECCV 2024 (Oral)

[arxiv] [project page[bibtex

Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Gedas Bertasius, ... , Michael Wray

CVPR 2024 (Oral)

[arxiv] [project website] [blog] [video] [bibtex

Learning To Recognize Procedural Activities with Distant Supervision

Xudong Lin, Fabio Petroni, Gedas Bertasius, Marcus Rohrbach, Shih-Fu Chang, Lorenzo Torresani

CVPR 2022

[arxiv] [code] [project page] [bibtex]

Strategic Video Intelligence

A central focus of our group is Strategic Video Intelligence, which asks models to perceive a dynamic scene, explain why it unfolds, simulate what could happen instead, and identify actions that could improve the outcome. We use team sports as our primary testbed. They combine complex multi-agent interaction with explicit rules and definitive outcomes, so inferences about goals, decisions, and consequences become testable.

Related Publications:

SVI-Bench: A Dynamic Microworld for Strategic Video Intelligence

Yulu PanHan YiSeongsu HaMd Mohaiminul IslamBenjamin ZhangLorenzo TorresaniGedas Bertasius

ECCV 2026

[arxiv] [video] [project page] [extended paper] [code] [data] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​

BASKET: A Large-Scale Video Dataset for Fine-Grained Skill Estimation

Yulu Pan, Ce Zhang, Gedas Bertasius

CVPR 2025

[arxiv] [project page] [code] [data] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

Egocentric Basketball Motion Planning from a Single First-Person Image
Gedas Bertasius, Aaron Chan and Jianbo Shi

CVPR 2018
[arxiv] [results] [MIT SSAC Poster] ​[bibtex]

Am I a Baller? Basketball Performance Assessment from First-Person Videos
Gedas Bertasius, Stella X. Yu, Hyun Soo Park and Jianbo Shi

​ICCV 2017
[​arxiv] [results] [bibtex

Video for Robotics

Robots that work alongside people must learn from them. We develop methods that translate observed human behavior into robot action, from behavior-grounded manipulation, where a robot watches a person act and carries out a related task, to learning action representations from human demonstrations at web scale, to robust long-horizon execution.

Related Publications:

WatchAct: A Benchmark for Behavior-Grounded Robot Manipulation

Baiqi Li, Ce Zhang, Yu Fang, Yue Yang, Shangzhe Li, Mingyu Ding, Gedas Bertasius

arXiv 2026

[arxiv] [project page] [code] [data] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​

LiLo-VLA: Compositional Long-Horizon Manipulation via Linked Object-Centric Policies

Yue YangShuo ChengYu FangHomanga BharadhwajMingyu Ding

Gedas BertasiusDaniel Szafir

arXiv 2026

[arxiv] [project page] [video] [code] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

BOSS: Benchmark for Observation Space Shift in Long-Horizon Task

Yue Yang, Linfeng Zhao, Mingyu Ding, Gedas Bertasius, Daniel Szafir

Robotics and Automation Letters (RA-L) 2025

[arxiv] [bibtex​​​​​​​​​​​​​​​

ReBot: Scaling Robot Learning with Real-to-Sim-to-Real Robotic Video Synthesis

Yu Fang, Yue Yang, Xinghao Zhu, Kaiyuan Zheng, Gedas Bertasius, Daniel Szafir, Mingyu Ding

IROS 2025

[arxiv] [project page] [code] [bibtex

Generative Video Modeling

Our group also builds generative video models for multimodal creation and editing, with applications that include video-to-music generation, audio-visual editing, and third-to-first person video translation.

Related Publications:

TeDiO: Temporal Diagonal Optimization for Training-Free Coherent Video Diffusion

Nurislam Tursynbek, Zhiqiang Lao, Heather Yu, Gedas Bertasius, Marc Niethammer

CVPR 2026 Workshop on Agentic AI for Visual Media

[arxiv] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

V2M-Zero: Zero-Pair Time-Aligned Video-to-Music Generation

Yan-Bo LinJonah CasebeerLong MaiAniruddha Mahapatra

Gedas BertasiusNicholas J. Bryan

arXiv 2026

[arxiv] [project page] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

Zero-Shot Audio-Visual Editing via Cross-Modal Delta Denoising

Yan-Bo Lin, Kevin Lin, Zhengyuan Yang, Linjie Li, Jianfeng Wang, Chung-Ching Lin, Xiaofei Wang, Gedas Bertasius, Lijuan Wang

WACV 2026 (Oral)

[arxiv] [project page] [code] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​​​​​​

VMAs: Video-to-Music Generation via Semantic Alignment in Web Music Videos

Yan-Bo Lin, Yu Tian, Linjie Yang, Gedas Bertasius, Heng Wang

WACV 2025 (Oral)

[arxiv] [project page] [code] [bibtex​​​​​​​​​​​​​​​​​​​​​​​​

4Diff: 3D-Aware Diffusion Model for Third-to-First Viewpoint Translation

Feng Cheng*, Mi Luo*, Huiyu Wang, Alex Dimakis, Lorenzo Torresani, Gedas Bertasius, Kristen Grauman

ECCV 2024

[arxiv] [bibtex

Contact

Prospective Graduate Students: I am recruiting motivated students in computer vision. Please email me a list of your prior publications and your CV.

​

Undergraduates at UNC: If you are interested in computer vision, especially its applications to sports, email me your CV and transcript with your GPA.

©2024 by Gedas Bertasius

bottom of page