A curated list of egocentric vision resources.
Egocentric (first-person) vision is a sub-field of computer vision that analyses image/video data obtained using a wearable camera simulating a person's visual field.
New to egocentric vision? A few landmark resources already in this list are a good entry point:
- Start here: An Outlook into the Future of Egocentric Vision (IJCV 2024) is a broad survey of the field's tasks, datasets, and open challenges
- Foundational datasets: Ego4D, EPIC-Kitchens 2020, Ego-Exo4D
Papers below are grouped by task first (see Papers), then cross-listed by venue for browsing recent conference proceedings; each section is collapsed by default, click "Show papers" to expand. Datasets are listed separately in Datasets, with a highlights table of flagship datasets followed by the full index.
-
Clustered into various problem statements.
- Action/Activity Recognition
- Object/Hand Recognition
- Action/Gaze Anticipation
- Localization
- Clustering
- Video Summarization
- Social Interactions
- Pose Estimation
- Human Object Interaction
- Temporal Boundary Detection
- Privacy in Egocentric Videos
- Multiple Egocentric Tasks
- Task Understanding
- Ego-Exo Cross-View Learning
- Egocentric Video-Language Models & Question Answering
- Egocentric Video Generation & World Models
- 3D Scene Reconstruction & Mapping
- Assistive & Navigation
- Miscellaneous (New Tasks)
Clustered according to the conferences.
Clustered in various problem statements.
Show papers (44)
-
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition - Sanjoy Kundu, Shanmukha Vellamcheti, and Sathyanarayanan N. Aakur. In ICCV 2025.
-
Understanding Multi-Task Activities from Single-Task Videos - Yuhan Shen and Ehsan Elhamifar. In CVPR 2025.
-
Test-Time Adaptation for Combating Missing Modalities in Egocentric Videos - Merey Ramazanova, Alejandro Pardo, Bernard Ghanem, and Motasem Alfarra. In ICLR 2025.
-
On the Utility of 3D Hand Poses for Action Recognition - Md Salman Shamil, Dibyadip Chatterjee, Fadime Sener, Shugao Ma, and Angela Yao. In ECCV 2024. [project page]
-
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos - Changan Chen, Kumar Ashutosh, Rohit Girdhar, David Harwath, and Kristen Grauman. In CVPR 2024. [project page]
-
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization - Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin, Eric Sauser, Bernt Schiele, and Shugao Ma. In CVPR 2024. [code]
-
Progress-Aware Online Action Segmentation for Egocentric Procedural Task Videos - Yuhan Shen and Ehsan Elhamifar. In CVPR 2024. [code]
-
TIM: A Time Interval Machine for Audio-Visual Action Recognition - Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zisserman, and Dima Damen. In CVPR 2024. [project page] [code]
-
Multimodal Distillation for Egocentric Action Recognition - Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens, and Tinne Tuytelaars. In ICCV 2023. [code]
-
What can a cook in Italy teach a mechanic in India? Action Recognition Generalisation Over Scenarios and Locations - Chiara Plizzari, Toby Perrett, Barbara Caputo, and Dima Damen. In ICCV 2023. [project page] [code]
-
MMG-Ego4D: Multimodal Generalization in Egocentric Action Recognition - Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, YILEI LI, Zhangyang Wang, Rakesh Ranjan. In CVPR 2023.
-
Therbligs In Action: Video Understanding through Motion Primitives - Eadom Dessalene, Michael Maynord, Cornelia Fermu ̈ller, Yiannis Aloimonos. In CVPR 2023. [project page]
-
[Learning Video Representations from Large Language Models](https://arxiv.org/pdf/2212.04501.pdf; https://facebookresearch.github.io/LaViLa) - Yue Zhao, Ishan Misra, Philipp Krähenbühl, Rohit Girdhar. In CVPR 2023. [project page] [code] [demo]
-
Learning State-Aware Visual Representations from Audible Interactions - Himangi Mittal, Pedro Morgado, Unnat Jain, Abhinav Gupta. In NeurIPS 2022. [Code] [Video]
-
Egocentric Activity Recognition and Localization on a 3D Map - Miao Liu, Lingni Ma, Kiran Somasundaram, Yin Li, Kristen Grauman, James M. Rehg, Chao Li. In ECCV 2022.
-
SOS! Self-supervised Learning Over Sets Of Handled Objects In Egocentric Action Recognition - Victor Escorcia, Ricardo Guerrero, Xiatian Zhu, Brais Martinez. In ECCV 2022.
-
E2(GO)MOTION: Motion Augmented Event Stream for Egocentric Action Recognition - Chiara Plizzari, Mirco Planamente, Gabriele Goletto, Marco Cannici, Emanuele Gusso, Matteo Matteucci, Barbara Caputo. In CVPR 2022.
-
Domain Generalization through Audio-Visual Relative Norm Alignment in First Person Action Recognition - Mirco Planamente, Chiara Plizzari, Emanuele Alberti, and Barbara Caputo. In WACV 2022.
-
With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition - Evangelos Kazakos, Jaesung Huh, Arsha Nagrani, Andrew Zisserman, and Dima Damen. In BMVC 2021. [project page] [code]
-
Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips - Lijin Yang, Yifei Huang, Yusuke Sugano, and Yoichi Sato. In BMVC 2021. [project page]
-
Interactive Prototype Learning for Egocentric Action Recognition - Xiaohan Wang, Linchao Zhu, Heng Wang, and Yi Yang. In ICCV 2021.
-
Multi-Modal Domain Adaptation for Fine-Grained Action Recognition - Jonathan Munro and Dima Damen. In CVPR 2020. [project page] [code]
-
Integrating Human Gaze Into Attention for Egocentric Activity Recognition - Kyle Min, Jason J. Corso. In WACV 2021. [code]
-
EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition - Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. In ICCV 2019. [code] [project page]
-
LSTA: Long Short-Term Attention for Egocentric Action Recognition - Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz. In CVPR 2019. [code]
-
Egocentric Activity Recognition on a Budget - Rafael Possas, Sheila Pinto Caceres, and Fabio Ramos. In CVPR 2018. [demo]
-
From Lifestyle VLOGs to Everyday Interaction - David F. Fouhey, Weicheng Kuo, Alexei A. Efros, and Jitendra Malik. In CVPR 2018. [project page]
-
Actor and Observer: Joint Modeling of First and Third-Person Videos - Gunnar A. Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. In CVPR 2018. [code]
-
In the eye of beholder: Joint learning of gaze and actions in first person video - Yin Li, Miao Liu, and James M. Rehg. In ECCV 2018.
-
Privacy-Preserving Human Activity Recognition from Extreme Low Resolution - Michael S. Ryoo, Brandon Rothrock, Charles Fleming, and Hyun Jong Yang. In AAAI 2017.
-
Jointly Recognizing Object Fluents and Tasks in Egocentric Videos - Yang Liu, Ping Wei, and Song-Chun Zhu. In ICCV 2017.
-
Trajectory Aligned Features For First Person Action Recognition - Suriya Singh, Chetan Arora, and C.V. Jawahar. In Pattern Recognition 2017.
-
First Person Action Recognition Using Deep Learned Descriptors - Suriya Singh, Chetan Arora, and C.V. Jawahar. In CVPR 2016. [project page] [code]
-
Understanding Hand-Object Manipulation with Grasp Types and Object Attributes - Minjie Cai, Kris M. Kitani, and Yoichi Sato. In Robotics: Science and Systems 2016.
-
Delving into egocentric actions - Yin Li, Zhefan Ye, and James M. Rehg. In CVPR 2015.
-
Pooled Motion Features for First-Person Videos - Michael S. Ryoo, Brandon Rothrock, and Larry H. Matthies. In CVPR 2015.
-
Generating Notifications for Missing Actions: Don't forget to turn the lights off! - Bilge Soran, Ali Farhadi, and Linda Shapiro. In ICCV 2015.
-
First-Person Activity Recognition: What Are They Doing to Me? - M. S. Ryoo and Larry Matthies. In CVPR 2013.
-
Detecting activities of daily living in first-person camera views - Hamed Pirsiavash and Deva Ramanan. In CVPR 2012.
-
Learning to recognize daily actions using gaze - Alireza Fathi, Yin Li, and James M. Rehg. In ECCV 2012.
-
Learning to recognize objects in egocentric activities - Alireza Fathi, Xiaofeng Ren, and James M. Rehg. In CVPR 2011.
-
Fast unsupervised ego-action learning for first-person sports videos - Kris M. Kitani, Takahiro Okabe, Yoichi Sato, and Akihiro Sugimoto. In CVPR 2011. [project page]
-
Temporal segmentation and activity classification from first-person sensing - Ekaterina H. Spriggs, Fernando De La Torre, and Martial Hebert. In CVPR Workshops 2009.
-
Wearable hand activity recognition for event summarization - W.W. Mayol and D.W. Murray. In IEEE International Symposium on Wearable Computers, 2005.
Show papers (24)
-
Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video - Yuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li, and Jingjing Zhang. In CVPR 2026.
-
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions - Taegyoon Yoon, Yegyu Han, Seojin Ji, Jaewoo Park, Sojeong Kim, Taein Kwon, and Hyung-Sin Kim. In CVPR 2026. [project page] [code]
-
Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention - Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu, Keze Wang, and Liang Lin. In NeurIPS 2025.
-
Is Tracking Really More Challenging in First Person Egocentric Vision? - Matteo Dunnhofer, Zaira Manigrasso, and Christian Micheloni. In ICCV 2025. [project page] [code]
-
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos - Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolandos Alexandros Potamias. In CVPR 2025. [project page]
-
HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos - Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, et al. In CVPR 2025. [project page]
-
ActionVOS: Actions as Prompts for Video Object Segmentation - Liangyang Ouyang, Ruicong Liu, Yifei Huang, Ryosuke Furuta, and Yoichi Sato. In ECCV 2024. [code]
-
Instance Tracking in 3D Scenes from Egocentric Videos - Yunhan Zhao, Haoyu Ma, Shu Kong, and Charless Fowlkes. In CVPR 2024. [code]
-
Learning to Segment Referred Objects from Narrated Egocentric Videos - Yuhan Shen, Huiyu Wang, Xitong Yang, Matt Feiszli, Ehsan Elhamifar, Lorenzo Torresani, and Effrosyni Mavroudi. In CVPR 2024.
-
EgoTracks: A Long-term Egocentric Visual Object Tracking Dataset - Hao Tang, Kevin J Liang, Kristen Grauman, Matt Feiszli, and Weiyao Wang. In NeurIPS 2023. [dataset]
-
Self-Supervised Object Detection from Egocentric Videos - Peri Akiva, Jing Huang, Kevin J Liang, Rama Kovvuri, Xingyu Chen, Matt Feiszli, Kristin Dana, and Tal Hassner. In ICCV 2023.
-
EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding - Chenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei, Jiabo Hu, Hichem El-Mohri, Sean Culatana, Roshan Sumbaly, and Zhicheng Yan. In ICCV 2023. [project page] [code]
-
Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB Videos - Yilin Wen, Hao Pan, Lei Yang, Jia Pan, Taku Komura, Wenping Wang. In CVPR 2023. [Code]
-
Generative Adversarial Network for Future Hand Segmentation from Egocentric Video - Wenqi Jia, Miao Liu, James M. Rehg. In ECCV 2022.
-
Whose Hand Is This? Person Identification From Egocentric Hand Gestures - Satoshi Tsutsui, Yanwei Fu, and David J. Crandall. In WACV 2021.
-
Generalizing Hand Segmentation in Egocentric Videos with Uncertainty-Guided Model Adaptation - Minjie Cai, Feng Lu, and Yoichi Sato. In CVPR 2020. [code]
-
H+O: Unified Egocentric Recognition of 3D Hand-Object Poses and Interactions - Bugra Tekin, Federica Bogo, and Marc Pollefeys. In CVPR 2019. [video]
-
Analysis of Hand Segmentation in the Wild - Aisha Urooj Khan and Ali Borji. In CVPR 2018.
-
First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations - Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. In CVPR 2018. [project page] [code]
-
Egocentric Gesture Recognition Using Recurrent 3D Convolutional Neural Networks with Spatiotemporal Transformer Modules - Congqi Cao, Yifan Zhang, Yi Wu, Hanqing Lu, and Jian Cheng. In ICCV 2017.
-
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions - Sven Bambach, Stefan Lee, David J. Crandall, and Chen Yu. In ICCV 2015.
-
Detecting Snap Points in Egocentric Video with a Web Photo Prior - Bo Xiong and Kristen Grauman. In ECCV 2014. [project page] [code]
-
Pixel-level hand detection in ego-centric videos - Cheng Li and Kris M. Kitani. In CVPR 2013. [video] [code]
-
Context-based vision system for place and object recognition - Antonio Torralba, Kevin P. Murphy, William T. Freeman, Mark A. Rubin. In ICCV 2003. [project page]
Show papers (24)
-
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span - Heeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock, and Gunhee Kim. In NeurIPS 2025.
-
HOIGaze: Gaze Estimation During Hand-Object Interactions in Extended Reality Exploiting Eye-Hand-Head Coordination - Zhiming Hu, Daniel Haeufle, Syn Schmitt, and Andreas Bulling. In SIGGRAPH 2025. [project page] [code]
-
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency - Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu, Fanman Meng, Lili Pan, and Hongliang Li. In CVPR 2026. [code]
-
Forecasting 3D Scanpaths in Egocentric Video - Fiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman, James M. Rehg, Vamsi Krishna Ithapu, and Calvin Murdock. In CVPR 2026.
-
FIction: 4D Future Interaction Prediction from Video - Kumar Ashutosh, Georgios Pavlakos, and Kristen Grauman. In CVPR 2025. [code]
-
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation - Bolin Lai, Fiona Ryan, Wenqi Jia, Miao Liu, and James M. Rehg. In ECCV 2024. [project page]
-
AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation - Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Jose J. Guerrero, Giovanni Maria Farinella, and Antonino Furnari. In ECCV 2024. [code]
-
PALM: Predicting Actions through Language Models - Sanghwan Kim, Daoji Huang, Yongqin Xian, Otmar Hilliges, Luc Van Gool, and Xi Wang. In ECCV 2024.
-
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation - Razvan-George Pasca, Alexey Gavryushin, Muhammad Hamza, Yen-Ling Kuo, Kaichun Mo, Luc Van Gool, Otmar Hilliges, and Xi Wang. In CVPR 2024. [project page]
-
Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting - Wentao Bao, Lele Chen, Libing Zeng, Zhong Li, Yi Xu, Junsong Yuan, and Yu Kong. In ICCV 2023. [project page] [code]
-
Intention-Conditioned Long-Term Human Egocentric Action Forecasting - Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee. In WACV 2023.
-
A Hybrid Egocentric Activity Anticipation Framework via Memory-Augmented Recurrent and One-shot Representation Forecasting - Tianshan Liu and Kin-Man Lam. In CVPR 2022.
-
Learning to Anticipate Egocentric Actions by Imagination - Yu Wu, Linchao Zhu, Xiaohan Wang, Yi Yang, and Fei Wu. In TIP 2021.
-
Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video - Miao Liu, Siyu Tang, Yin Li, and James M. Rehg. In ECCV 2020. [project page]
-
How Can I See My Future? FvTraj: Using First-person View for Pedestrian Trajectory Prediction - Huikun Bi, Ruisi Zhang, Tianlu Mao, Zhigang Deng, and Zhaoqi Wang. In ECCV 2020. [presentation video] [summary video]
-
Multimodal Future Localization and Emergence Prediction for Objects in Egocentric View With a Reachability Prior - Osama Makansi, Ozgun Cicek, Kevin Buchicchio, and Thomas Brox. In CVPR 2020. [demo] [code] [project page]
-
EGO-TOPO: Environment Affordances from Egocentric Video - Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman. In CVPR 2020. [project page] [demo]
-
What Would You Expect? Anticipating Egocentric Actions with Rolling-Unrolling LSTMs and Modality Attention - Antonino Furnari and Giovanni Maria Farinella. In ICCV 2019 [code] [demo]
-
Digging Deeper into Egocentric Gaze Prediction - Hamed R. Tavakoli, Esa Rahtu, Juho Kannala, and Ali Borji. In WACV 2019.
-
Predicting Gaze in Egocentric Video by Learning Task-dependent Attention Transition - Yifei Huang, Minjie Cai, Zhenqiang Li, and Yoichi Sato. In ECCV 2018 [code]
-
First-Person Activity Forecasting with Online Inverse Reinforcement Learning - Nicholas Rhinehart and Kris M. Kitani. In ICCV 2017. [video]
-
Deep future gaze: Gaze anticipation on egocentric videos using adversarial networks - Mengmi Zhang, Keng Teck Ma, Joo Hwee Lim, Qi Zhao, and Jiashi Feng. In CVPR 2017. [code]
-
Going deeper into first-person activity recognition - Minghuang Ma, Haoqi Fan, and Kris M. Kitani. In CVPR 2016.
-
Learning to predict gaze in egocentric video - Yin Li, Alireza Fathi, and James M. Rehg. In ICCV 2013.
Show papers (13)
-
Beyond Caption-Based Queries for Video Moment Retrieval - David Pujol-Perich, Albert Clapés, Dima Damen, Sergio Escalera, and Michael Wray. In CVPR 2026.
-
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization - Bing Fan, Yunhe Feng, Yapeng Tian, James Chenhao Liang, Yuewei Lin, Yan Huang, and Heng Fan. In ICCV 2025. [code]
-
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance - Mingfang Zhang, Ryo Yonetani, Yifei Huang, Liangyang Ouyang, Ruicong Liu, and Yoichi Sato. In ICCV 2025.
-
Spatial Cognition from Egocentric Video: Out of Sight, Not Out of Mind - Chiara Plizzari, Shubham Goel, Toby Perrett, Jacob Chalk, Angjoo Kanazawa, and Dima Damen. In 3DV 2025. [project page]
-
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory - Zaira Manigrasso, Matteo Dunnhofer, Antonino Furnari, Moritz Nottebaum, Antonio Finocchiaro, Davide Marana, Rosario Forte, Giovanni Maria Farinella, and Christian Micheloni. In WACV 2026.
-
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos - Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar, Jacob Donley, Chao Li, Gunhee Kim, Vamsi Krishna Ithapu, and Calvin Murdock. In ECCV 2024. [project page]
-
EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries - Jinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao, and Bernard Ghanem. In ICCV 2023. [code]
-
Hand-Priming in Object Localization for Assistive Egocentric Vision - Kyungjun Lee, Abhinav Shrivastava, and Hernisa Kacorri. In WACV 2020.
-
Egocentric Shopping Cart Localization - Emiliano Spera, Antonino Furnari, Sebastiano Battiato, and Giovanni Maria Farinella. In ICPR 2018.
-
Recognizing personal locations from egocentric videos - Antonino Furnari, Giovanni Maria Farinella, and Sebastiano Battiato. In IEEE Transactions on Human-Machine Systems 2017.
-
Personal-Location-Based Temporal Segmentation of Egocentric Video for Lifelogging Applications - Antonino Furnari, Sebastiano Battiato, and Giovanni Maria Farinella. In Journal of Visual Communication and Image Representation 2017. [demo] [project page]
-
Egocentric Future Localization - Hyun Soo Park, Jyh-Jing Hwang, Yedong Niu, and Jianbo Shi. In CVPR 2016. [demo]
-
Real-time localization and mapping with wearable active vision - A.J. Davison, W.W. Mayol, and D.W. Murray. In The Second IEEE and ACM International Symposium 2003.
Show papers (2)
-
Sr-clustering: Semantic regularized clustering for egocentric photo streams segmentation - Mariella Dimiccoli, Marc Bolanosa, Estefania Talavera Maedeh Aghaei, Stavri G. Nikolov, and Petia Radeva. In Computer Vision and Image Understanding 2017.
-
Summarization and Classification of Wearable Camera Streams by Learning the Distributions over Deep Features of Out-of-Sample Image Sequences - Alessandro Perina, Sadegh Mohammadi, Nebojsa Jojic, and Vittorio Murino. In ICCV 2017.
Show papers (4)
-
Query-focused video summarization: Dataset, evaluation, and a memory network based approach - Aidean Sharghi, Jacob S. Laurel and Boqing Gong. In CVPR 2017.
-
Toward storytelling from visual lifelogging: An overview - Marc Bolanos, Mariella Dimiccoli, and Petia Radeva. In IEEE Transactions on Human-Machine Systems 2017.
-
Story-Driven Summarization for Egocentric Video - Zheng Lu and Kristen Grauman. In CVPR 2013 [project page]
-
Discovering Important People and Objects for Egocentric Video Summarization - Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman. In CVPR 2012. [project page]
Show papers (6)
-
Seeing Conversations: Communication Context Identification in Egocentric Video - Tobias Dorszewski and Jens Hjortkjær. In CVPR 2026.
-
Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding - Minh Tran, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Min Sun, and Mohammad Soleymani. In ECCV 2024.
-
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective - Wenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla, James M. Rehg, Vamsi Krishna Ithapu, and Ruohan Gao. In CVPR 2024. [project page]
-
EgoCom: A Multi-person Multi-modal Egocentric Communications Dataset - Curtis G. Northcutt, Shengxin Zha, Steven Lovegrove, and Richard Newcombe. In PAMI 2020.
-
Deep Dual Relation Modeling for Egocentric Interaction Recognition - Haoxin Li, Yijun Cai, and Wei-Shi Zheng. In CVPR 2019.
-
Recognizing Micro-Actions and Reactions from Paired Egocentric Videos - Ryo Yonetani, Kris M. Kitani, and Yoichi Sato. In CVPR 2016.
Show papers (50)
-
EgoHumans: An Egocentric 3D Multi-Human Benchmark - Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard Newcombe, Minh Vo, and Kris Kitani. In ICCV 2023 (Oral). [code]
-
Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation - Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Jiahe Zhao, Michael J. Black, and Dima Damen. In ECCV 2026. [project page] [code]
-
E-3DPSM: A State Machine for Event-Based Egocentric 3D Human Pose Estimation - Mayur Deshmukh, Hiroyasu Akada, Helge Rhodin, Christian Theobalt, and Vladislav Golyanik. In CVPR 2026. [project page]
-
Egocentric Visibility-Aware Human Pose Estimation - Peng Dai, Yu Zhang, Yiqiang Feng, Zhen Fan, and Yang Zhang. In CVPR 2026.
-
EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR - Zhenyu Li, Sai Kumar Dwivedi, Filip Maric, Carlos Chacon, Nadine Bertsch, Filippo Arcadu, et al. In CVPR 2026. [project page]
-
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains - Wiktor Mucha, Michael Wray, and Martin Kampel. In WACV 2026.
-
UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation - Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka, Juan Carlos Niebles, and Ehsan Adeli. In ICCV 2025. [project page] [code]
-
Head2Body: Body Pose Generation from Multi-sensory Head-mounted Inputs - Minh Tran, Hongda Mao, Qingshuang Chen, and Yelin Kim. In ICCV 2025.
-
Bring Your Rear Cameras for Egocentric 3D Human Pose Estimation - Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. In ICCV 2025. [project page]
-
Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric Vision - Tianma Shen, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar, Ryan Fell, Julianna Dietrich, Maria Kyrarini, Christopher Kitts, and David C. Jeong. In ICCV 2025. [project page]
-
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba - Quang Nguyen, Nhat Le, Baoru Huang, Minh Nhat Vu, Chengcheng Tang, Van Nguyen, Ngan Le, Thieu Vo, and Anh Nguyen. In ICCV 2025.
-
REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning - Jihyun Lee, Weipeng Xu, Alexander Richard, Shih-En Wei, Shunsuke Saito, Shaojie Bai, Te-Li Wang, Minhyuk Sung, Tae-Kyun Kim, and Jason Saragih. In CVPR 2025. [project page]
-
FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video - Andrea Boscolo Camiletto, Jian Wang, Eduardo Alvarado, Rishabh Dabral, Thabo Beeler, Marc Habermann, and Christian Theobalt. In CVPR 2025. [project page] [code]
-
EgoLM: Multi-Modal Language Model of Egocentric Motions - Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim, Yuting Ye, Richard Newcombe, Ziwei Liu, and Lingni Ma. In CVPR 2025. [project page]
-
Estimating Body and Hand Motion in an Ego-sensed World - Brent Yi, Vickie Ye, Maya Zheng, Yunqi Li, Lea Müller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo Kanazawa. In CVPR 2025. [project page]
-
Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera - Zhengdi Yu, Stefanos Zafeiriou, and Tolga Birdal. In CVPR 2025. [project page]
-
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision - Yiming Zhao, Taein Kwon, Paul Streli, Marc Pollefeys, and Christian Holz. In CVPR 2025. [project page]
-
Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input - Jian Wang, Rishabh Dabral, Diogo Luvizon, Zhe Cao, Lingjie Liu, Thabo Beeler, and Christian Theobalt. In CVPR 2025. [project page]
-
EgoCast: Forecasting Egocentric Human Pose in the Wild - Maria Escobar, Juanita Puentes, Cristhian Forigua, Jordi Pont-Tuset, Kevis-Kokitsi Maninis, and Pablo Arbelaez. In WACV 2025. [code]
-
Social EgoMesh Estimation - Luca Scofano, Alessio Sampieri, Edoardo De Matteis, Indro Spinelli, and Fabio Galasso. In WACV 2025. [code]
-
Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data - Seunggeun Chi, Pin-Hao Huang, Enna Sachdeva, Hengbo Ma, Karthik Ramani, and Kwonjoon Lee. In NeurIPS 2024. [project page]
-
EgoSim: An Egocentric Multi-view Simulator and Real Dataset for Body-worn Cameras during Motion and Activity - Dominik Hollidt, Paul Streli, Jiaxi Jiang, Yasaman Haghighi, Changlin Qian, Xintong Liu, and Christian Holz. In NeurIPS 2024. [project page]
-
Nymeria: A Massive Collection of Egocentric Multi-modal Human Motion in the Wild - Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, et al. In ECCV 2024. [project page]
-
EgoPoseFormer: A Simple Baseline for Stereo Egocentric 3D Human Pose Estimation - Chenhongyi Yang, Anastasia Tkach, Shreyas Hampali, Linguang Zhang, Elliot J. Crowley, and Cem Keskin. In ECCV 2024. [code]
-
EgoPoser: Robust Real-Time Egocentric Pose Estimation from Sparse and Intermittent Observations Everywhere - Jiaxi Jiang, Paul Streli, Manuel Meier, and Christian Holz. In ECCV 2024. [project page]
-
3D Hand Pose Estimation in Everyday Egocentric Images - Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. In ECCV 2024. [project page]
-
EgoBody3M: Egocentric Body Tracking on a VR Headset using a Diverse Dataset - Amy Zhao, Chengcheng Tang, Lezi Wang, Yijing Li, Mihika Dave, Lingling Tao, Christopher D. Twigg, and Robert Y. Wang. In ECCV 2024. [dataset]
-
Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects - Zicong Fan, Takehiko Ohkawa, Linlin Yang, Nie Lin, Zhishan Zhou, Shihao Zhou, et al. In ECCV 2024.
-
EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams - Christen Millerdurai, Hiroyasu Akada, Jian Wang, Diogo Luvizon, Christian Theobalt, and Vladislav Golyanik. In CVPR 2024. [project page]
-
3D Human Pose Perception from Egocentric Stereo Videos - Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. In CVPR 2024. [code]
-
Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement - Jian Wang, Zhe Cao, Diogo Luvizon, Lingjie Liu, Kripasindhu Sarkar, Danhang Tang, Thabo Beeler, and Christian Theobalt. In CVPR 2024. [code]
-
Attention-Propagation Network for Egocentric Heatmap to 3D Pose Lifting - Taeho Kang and Youngki Lee. In CVPR 2024. [code]
-
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation - Ruicong Liu, Takehiko Ohkawa, Mingfang Zhang, and Yoichi Sato. In CVPR 2024. [code]
-
Real-Time Simulated Avatar from Head-Mounted Sensors - Zhengyi Luo, Jinkun Cao, Rawal Khirodkar, Alexander Winkler, Jing Huang, Kris Kitani, and Weipeng Xu. In CVPR 2024. [project page]
-
Mocap Everyone Everywhere: Lightweight Motion Capture With Smartwatches and a Head-Mounted Camera - Jiye Lee and Hanbyul Joo. In CVPR 2024. [project page]
-
Spectral Graphormer: Spectral Graph-Based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images - Tze Ho Elden Tse, Franziska Mueller, Zhengyang Shen, Danhang Tang, Thabo Beeler, Mingsong Dou, Yinda Zhang, Sasa Petrovic, Hyung Jin Chang, Jonathan Taylor, and Bardia Doosti. In ICCV 2023. [project page]
-
Probabilistic Human Mesh Recovery in 3D Scenes from Egocentric Views - Siwei Zhang, Qianli Ma, Yan Zhang, Sadegh Aliakbarian, Darren Cosker, and Siyu Tang. In ICCV 2023. [project page] [code]
-
AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose Estimation - Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, LUAN TRAN, Cem Keskin. In CVPR 2023.
-
Scene-aware Egocentric 3D Human Pose Estimation - Jian Wang, Diogo Luvizon, Weipeng Xu, Lingjie Liu, Kripasindhu Sarkar, Christian Theobalt. In CVPR 2023.
-
Ego-Body Pose Estimation via Ego-Head Pose Estimation - Jiaman Li · Karen Liu · Jiajun Wu. In CVPR 2023.
-
EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices - Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, Siyu Tang. In ECCV 2022. [project page] [dataset] [code]
-
UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture - Hiroyasu Akada, Jian Wang, Soshi Shimada, Masaki Takahashi, Christian Theobalt, Vladislav Golyanik. In ECCV 2022. [project page] [code] [dataset] [demo]
-
Estimating Egocentric 3D Human Pose in the Wild with External Weak Supervision - Jian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Diogo Luvizon, Christian Theobalt. In CVPR 2022. [project page]
-
Estimating Egocentric 3D Human Pose in Global Space - Jian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Christian Theobalt. In ICCV 2021. [project page]
-
Automatic Calibration of the Fisheye Camera for Egocentric 3D Human Pose Estimation From a Single Image - Yahui Zhang, Shaodi You, and Theo Gevers. In WACV 2021.
-
You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions - Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman. In CVPR 2020. [demo] [project page] [dataset] [code]
-
Ego-Pose Estimation and Forecasting as Real-Time PD Control - Ye Yuan and Kris Kitani. In ICCV 2019. [code] [project page] [demo]
-
xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera - Denis Tome, Patrick Peluse, Lourdes Agapito, and Hernan Badino. In ICCV 2019. [demo] [dataset]
-
Seeing Invisible Poses: Estimating 3D Body Pose from Egocentric Video - Hao Jiang and Kristen Grauman. In CVPR 2017.
-
First-Person Pose Recognition using Egocentric Workspaces - Gregory Rogez, James S. Supancic, and Deva Ramanan. In CVPR 2015.
Show papers (20)
-
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation - Bohan Zhou, Yi Zhan, Zhongbin Zhang, and Zongqing Lu. In NeurIPS 2025. [project page] [code]
-
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions - Liang Xu, Chengqun Yang, Zili Lin, Fei Xu, Yifan Liu, Congsheng Xu, et al. In ICCV 2025. [project page]
-
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation - Gen Li, Nikolaos Tsagkas, Jifei Song, Ruaridh Mon-Williams, Sethu Vijayakumar, Kun Shao, and Laura Sevilla-Lara. In ICCV 2025. [project page]
-
ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos - Yuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang, Zhaojie Fang, Chenghong Li, and Xiaoguang Han. In CVPR 2026. [project page] [code]
-
EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation - Abhishek Saroha, Huajian Zeng, Xingxing Zuo, Daniel Cremers, and Xi Wang. In CVPR 2026. [project page] [code]
-
ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions - Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. In CVPR 2025. [code]
-
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision - Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, and Shinsuke Mori. In CVPR 2025. [project page] [code]
-
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction - Yuejiao Su, Yi Wang, Qiongyang Hu, Chuang Yang, and Lap-Pui Chau. In CVPR 2025. [project page]
-
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views - Yuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu, Yang Cao, and Zheng-Jun Zha. In NeurIPS 2024. [project page] [code]
-
Are Synthetic Data Useful for Egocentric Hand-Object Interaction Detection? - Rosario Leonardi, Antonino Furnari, Francesco Ragusa, and Giovanni Maria Farinella. In ECCV 2024. [project page] [code]
-
Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos - Zecheng Yu, Yifei Huang, Ryosuke Furuta, Takuma Yagi, Yusuke Goutsu, and Yoichi Sato. In WACV 2023.
-
EgoPCA: A New Framework for Egocentric Hand-Object Interaction Understanding - Yue Xu, Yong-Lu Li, Zhemin Huang, Michael Xu Liu, Cewu Lu, Yu-Wing Tai, and Chi-Keung Tang. In ICCV 2023. [project page]
-
ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation - Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, Otmar Hilliges. In CVPR 2023. [code]
-
Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications - Lingzhi Zhang, Shenghao Zhou, Simon Stent, Jianbo Shi. In ECCV 2022. [project page] [code] [dataset]
-
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction - Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, Li Yi. In CVPR 2022. [project page] [video]
-
Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction - Takuma Yagi, Md Tasnimul Hasan, and Yoichi Sato. In BMVC 2021. [project page] [code]
-
The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain - Francesco Ragusa, Antonino Furnari, Salvatore Livatino, and Giovanni Maria Farinella. In WACV 2021. [project page]
-
Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video - Miao Liu, Siyu Tang, Yin Li, and James M. Rehg. In ECCV 2020. [project page]
-
You-Do, I-Learn: Discovering Task Relevant Objects and their Modes of Interaction from Multi-User Egocentric Video - Dima Damen, Tessid Leelasawassuk, Osian Haines, Andrew Calway,and Walterio Mayol-Cuevas. In BMVC 2014 [project page]
-
Automated capture and delivery of assistive task guidance with an eyewear computer: the GlaciAR system - Teesid Leelasawassuk, Dima Damen, and Walterio Mayol-Cuevas. In Augmented Human International Conference, ACM 2017.
Show papers (4)
-
Streaming Detection of Queried Event Start - Cristóbal Eyzaguirre, Eric Tang, Shyamal Buch, Adrien Gaidon, Jiajun Wu, and Juan Carlos Niebles. In NeurIPS 2024. [project page]
-
Ego-Only: Egocentric Action Detection without Exocentric Transferring - Huiyu Wang, Mitesh Kumar Singh, and Lorenzo Torresani. In ICCV 2023.
-
Trespassing the Boundaries: Labeling Temporal Bounds for Object Interactions in Egocentric Video - Davide Moltisanti, Michael Wray, Walterio Mayol-Cuevas, and Dima Damen. In ICCV 2017.
-
Temporal segmentation of egocentric videos -Yair Poleg, Chetan Arora, and Shmuel Peleg. In CVPR 2014.
Show papers (4)
-
Is Sharing of Egocentric Video Giving Away Your Biometric Signature? - Daksh Thapar, Chetan Arora, and Aditya Nigam. In ECCV 2020. [project page]
-
Mitigating Bystander Privacy Concerns in Egocentric Activity Recognition with Deep Learning and Intentional Image Degradation - Mariella Dimiccoli, Juan Marin, and Edison Thomaz. In Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies 2018.
-
Privacy-Preserving Human Activity Recognition from Extreme Low Resolution - Michael S. Ryoo, Brandon Rothrock, Charles Fleming, and Hyun Jong Yang. In AAAI 2017.
-
Ego-Surfing First Person Videos - Ryo Yonetani, Kris M. Kitani, and Yoichi Sato. In CVPR 2015.
Show papers (12)
-
EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models - Andy Bonnetto, Haozhe Qi, Franklin Leong, Matea Tashkovska, Mahdi Rad, Solaiman Shokur, Friedhelm Hummel, Silvestro Micera, Marc Pollefeys, and Alexander Mathis. In NeurIPS 2025. [project page] [code]
-
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception - Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan, Calvin Murdock, Ishwarya Ananthabhotla, Yijun Qian, Vamsi Krishna Ithapu, Dinesh Manocha, and Ruohan Gao. In ICCV 2025. [project page]
-
EgoM2P: Egocentric Multimodal Multitask Pretraining - Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao, Marc Pollefeys, and Siyu Tang. In ICCV 2025. [project page]
-
HD-EPIC: A Highly-Detailed Egocentric Video Dataset - Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Parida, et al. In CVPR 2025. [project page]
-
Ego-VPA: Egocentric Video Understanding with Parameter-Efficient Adaptation - Tz-Ying Wu, Kyle Min, Subarna Tripathi, and Nuno Vasconcelos. In WACV 2025. [code]
-
A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives - Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, and Giuseppe Averta. In CVPR 2024.
-
An Outlook into the Future of Egocentric Vision - Chiara Plizzari, Gabriele Goletto, Antonino Furnari, Siddhant Bansal, Francesco Ragusa, Giovanni Maria Farinella, Dima Damen, and Tatiana Tommasi. In IJCV 2024. [project page]
-
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives - Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zach Chavis, Joya Chen, Feng Cheng, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong, Maria Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang, Md Mohaiminul Islam, Suyog Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Miguel Martin, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh Kumar Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Edward Zhang, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina Gonzalez, Prince Gupta, Jiabo Hu, Yifei Huang, Yiming Huang, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu, Mi Luo, Zhengyi Luo, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Wei Shan, Kiran Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Huiyu Wang, Yuchen Wang, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao, Ziwei Zhao, Zhifan Zhu, Jeff Zhuo, Pablo Arbelaez, Gedas Bertasius, David Crandall, Dima Damen, Jakob Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato, Manolis Savva, Jianbo Shi, Mike Zheng Shou, and Michael Wray. In CVPR 2024. [project page]
-
EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone - Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang. In ICCV 2023. [project page] [code]
-
Egocentric Video-Language Pretraining - Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu and Mike Zheng Shou. In NeurIPS 2022. [project page] [code]
-
Ego4D: Around the World in 3,000 Hours of Egocentric Video - Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Christian Fuegen, Abrham Gebreselasie, Cristina Gonzalez, James Hillis, Xuhua Huang, Yifei Huang, Wenqi Jia, Weslie Khoo, Jachym Kolar, Satwik Kottur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Yunyi Zhu, Pablo Arbelaez, David Crandall, Dima Damen, Giovanni Maria Farinella, Bernard Ghanem, Vamsi Krishna Ithapu, C.V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, and Jitendra Malik. In CVPR 2022. [Github] [project page] [video]
-
Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities - Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. In CVPR 2022. [project page]
Show papers (12)
-
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos - Yayuan Li, Aadit Jain, Filippos Bellos, and Jason J. Corso. In CVPR 2026. [project page]
-
EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios - Lu Qiu, Yi Chen, Yuying Ge, Yixiao Ge, Ying Shan, and Xihui Liu. In IJCV 2026. [project page]
-
EgoPlan-Bench: Benchmarking Multimodal Large Language Models for Human-Level Planning - Yi Chen, Yuying Ge, Yixiao Ge, Mingyu Ding, Bohao Li, Rui Wang, Ruifeng Xu, Ying Shan, and Xihui Liu. In IJCV 2026. [project page]
-
IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants - Vivek Chavan, Yasmina Imgrund, Tung Dao, Sanwantri Bai, Bosong Wang, Ze Lu, Oliver Heimann, and Jörg Krüger. In NeurIPS 2025. [project page] [code] [dataset]
-
HiERO: Understanding the Hierarchy of Human Behavior Enhances Reasoning on Egocentric Videos - Simone Alberto Peirone, Francesca Pistilli, and Giuseppe Averta. In ICCV 2025. [project page] [code]
-
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities - Michele Mazzamuto, Antonino Furnari, Yoichi Sato, and Giovanni Maria Farinella. In CVPR 2025.
-
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities - Rohith Peddi, Shivvrat Arya, Bharath Challa, Likhitha Pallapothula, Akshay Vyas, Bhavya Gouripeddi, et al. In NeurIPS 2024. [project page]
-
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos - Luigi Seminara, Giovanni Maria Farinella, and Antonino Furnari. In NeurIPS 2024. [code]
-
Error Detection in Egocentric Procedural Task Videos - Shih-Po Lee, Zijia Lu, Zekun Zhang, Minh Hoai, and Ehsan Elhamifar. In CVPR 2024. [project page] [code]
-
EgoTV: Egocentric Task Verification from Natural Language Task Descriptions - Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra, and Ruta Desai. In ICCV 2023. [project page] [code]
-
My View is the Best View: Procedure Learning from Egocentric Videos - Siddhant Bansal, Chetan Arora, C.V. Jawahar. In ECCV 2022. [project page] [dataset] [code]
-
AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant - Benita Wong, Joya Chen, You Wu, Stan Weixian Lei, Dongxing Mao, Difei Gao, Mike Zheng Shou. In ECCV 2022. [project page] [code]
Show papers (19)
-
Exo2Ego: Exocentric Knowledge Guided MLLM for Egocentric Video Understanding - Haoyu Zhang, Qiaohui Chu, Meng Liu, Haoxiang Shi, Yaowei Wang, and Liqiang Nie. In AAAI 2026. [project page]
-
SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion - Xiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu, Fanman Meng, Linfeng Xu, and Hongliang Li. In CVPR 2026. [code]
-
RegionAligner: Bridging Ego-Exo Views for Object Correspondence via Unified Text-Visual Learning - Yuhao Su and Ehsan Elhamifar. In WACV 2026.
-
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding - Ege Özsoy, Arda Mamur, Felix Tristram, Chantal Pellegrini, Magdalena Wysocki, Benjamin Busam, and Nassir Navab. In NeurIPS 2025. [code]
-
Robust Ego-Exo Correspondence with Long-Term Memory - Yijun Hu, Bing Fan, Xin Gu, Haiqing Ren, Dongfang Liu, Heng Fan, and Libo Zhang. In NeurIPS 2025. [code]
-
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs - Yuping He, Yifei Huang, Guo Chen, Baoqi Pei, Jilan Xu, Tong Lu, Jiangmiao Pang, et al. In NeurIPS 2025. [code]
-
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives - Yuqian Fu, Runze Wang, Bin Ren, Guolei Sun, Biao Gong, Yanwei Fu, Danda Pani Paudel, Xuanjing Huang, and Luc Van Gool. In ICCV 2025. [project page]
-
O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views - Lorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Perez-Yus, Ruben Martinez-Cantin, and Jose J. Guerrero. In ICCV 2025. [project page] [code]
-
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations - Jungin Park, Jiyoung Lee, and Kwanghoon Sohn. In CVPR 2025. [code]
-
Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning - Mi Luo, Zihui Xue, Alex Dimakis, and Kristen Grauman. In CVPR 2025. [project page]
-
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos - Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan, and Kristen Grauman. In CVPR 2025. [project page]
-
Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues - Sihong Huang, Jiaxin Wu, Xiaoyong Wei, Yi Cai, Dongmei Jiang, and Yaowei Wang. In CVPR 2025. [code]
-
EgoExo-Gen: Ego-centric Video Prediction by Watching Exo-centric Videos - Jilan Xu, Yifei Huang, Baoqi Pei, Junlin Hou, Qingqiu Li, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie. In ICLR 2025.
-
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos - Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura, Ryosuke Furuta, Atsushi Hashimoto, Yoshitaka Ushiku, and Yoichi Sato. In WACV 2025. [code]
-
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding - Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng, Jing-Ke Meng, and Wei-Shi Zheng. In ECCV 2024. [code]
-
4Diff: 3D-Aware Diffusion Model for Third-to-First Viewpoint Translation - Feng Cheng, Mi Luo, Huiyu Wang, Alex Dimakis, Lorenzo Torresani, Gedas Bertasius, and Kristen Grauman. In ECCV 2024. [project page]
-
Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs - Camillo Quattrocchi, Antonino Furnari, Daniele Di Mauro, Mario Valerio Giuffrida, and Giovanni Maria Farinella. In ECCV 2024. [code]
-
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World - Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, et al. In CVPR 2024. [code]
-
Fusing Personal and Environmental Cues for Identification and Segmentation of First-Person Camera Wearers in Third-Person Views - Ziwei Zhao, Yuchen Wang, Chuhua Wang, and David Crandall. In CVPR 2024. [code]
Show papers (39)
-
X-LeBench: A Benchmark for Extremely Long Egocentric Video Understanding - Wenqi Zhou, Kai Cao, Hao Zheng, Yunze Liu, Xinyi Zheng, Miao Liu, Per Ola Kristensson, Walterio W. Mayol-Cuevas, Fan Zhang, Weizhe Lin, and Junxiao Shen. In Findings of EMNLP 2025. [code]
-
Vinci: A Real-time Smart Assistant Based on Egocentric Vision-Language Model for Portable Devices - Yifei Huang, Jilan Xu, Baoqi Pei, Lijin Yang, Mingfang Zhang, Yuping He, Guo Chen, Xinyuan Chen, Yaohui Wang, Zheng Nie, Jinyao Liu, Dechen Lin, Fang Fang, Kunpeng Li, Chang Yuan, Yu Qiao, Yali Wang, and Limin Wang. In IMWUT 2025. [code]
-
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos - Shoubin Yu, Lei Shu, Antoine Yang, Yao Fu, Srinivas Sunkara, Maria Wang, Jindong Chen, Mohit Bansal, and Boqing Gong. In CVPR 2026. [project page] [code]
-
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics - Masatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato, and Takuma Yagi. In CVPR 2026. [project page]
-
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy - Jinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan, Yifan Yu, Dong Wang, Honglei Yan, Liang Yue, Shaofei Wang, Yixin Chen, Siyuan Huang, and Miao Liu. In CVPR 2026.
-
Ego-Grounding for Personalized Question-Answering in Egocentric Videos - Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, and Angela Yao. In CVPR 2026. [code]
-
EgoAVU: Egocentric Audio-Visual Understanding - Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja, Ernie Chang, Gregory P. Meyer, Gael Le Lan, Yunyang Xiong, Vikas Chandra, Yangyang Shi, Dinesh Manocha, and Zhipeng Cai. In CVPR 2026. [project page] [code]
-
EgoSound: Benchmarking Sound Understanding in Egocentric Videos - Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun, Tianwen Qian, Yuzheng Wu, Danda Pani Paudel, Xiangyang Xue, and Yanwei Fu. In CVPR 2026. [project page] [dataset]
-
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding - Arsha Nagrani, Jasper Uijlings, Shyamal Buch, Tobias Weyand, Sudheendra Vijayanarasimhan, Bo Hu, Ramin Mehran, David A Ross, and Cordelia Schmid. In CVPR 2026. [dataset]
-
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering - Yura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi, Jiankang Deng, and Stefanos Zafeiriou. In CVPR 2026. [project page]
-
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos - Peiran Wu, Yunze Liu, Miao Liu, and Junxiao Shen. In WACV 2026.
-
Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance - Francesco Ragusa, Michele Mazzamuto, Rosario Forte, Irene D'Ambra, James Fort, Jakob Engel, Antonino Furnari, and Giovanni Maria Farinella. In WACV 2026. [project page]
-
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining - Boshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng, and Qin Jin. In NeurIPS 2025. [code]
-
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT - Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He, Guo Chen, Fei Wu, Yu Qiao, and Jiangmiao Pang. In NeurIPS 2025. [code]
-
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? - Yuqian Yuan, Ronghao Dang, Long Li, Wentong Li, Dian Jiao, Xin Li, Deli Zhao, Fan Wang, Wenqiao Zhang, Jun Xiao, and Yueting Zhuang. In NeurIPS 2025. [project page] [code]
-
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios - Eun Chang, Zhuangqun Huang, Yiwei Liao, Sagar Ravi Bhavsar, Amogh Param, Tammy Stark, et al. In NeurIPS 2025.
-
Gaze-VLM: Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding - Anupam Pani and Yanchao Yang. In NeurIPS 2025. [code]
-
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting - Taiying Peng, Jiacheng Hua, Miao Liu, and Feng Lu. In NeurIPS 2025.
-
OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and Data - Hao Luo, Zihao Yue, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, Deheng Ye, and Zongqing Lu. In NeurIPS 2025. [code]
-
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video - Yulin Zhang, Cheng Shi, Yang Wang, and Sibei Yang. In NeurIPS 2025. [code]
-
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding - Yue Fan, Xiaojian Ma, Rongpeng Su, Jun Guo, Rujie Wu, Xi Chen, and Qing Li. In ICCV 2025.
-
Visual Intention Grounding for Egocentric Assistants - Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li, Arjun Akula, and Angela Yao. In ICCV 2025. [code]
-
EgoLife: Towards Egocentric Life Assistant - Jingkang Yang, Shuai Liu, Hongming Guo, et al. In CVPR 2025. [code]
-
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering - Sheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li, Xun Yang, Dan Guo, Meng Wang, Tat-Seng Chua, and Angela Yao. In CVPR 2025. [code]
-
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos - Chiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha, and Federico Tombari. In CVPR 2025. [code]
-
Object-Shot Enhanced Grounding Network for Egocentric Video - Yisen Feng, Haoyu Zhang, Meng Liu, Weili Guan, and Liqiang Nie. In CVPR 2025. [code]
-
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark - Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin, Boqiang Zhang, Long Li, Liuyi Wang, Qinyang Zeng, Xin Li, and Lidong Bing. In CVPR 2025. [code]
-
MM-Ego: Towards Building Egocentric Multimodal LLMs for Video QA - Hanrong Ye, Haotian Zhang, Erik Daxberger, Lin Chen, Zongyu Lin, et al. In ICLR 2025. [project page]
-
Do Egocentric Video-Language Models Truly Understand Hand-Object Interactions? - Boshen Xu, Ziheng Wang, Yang Du, Zhinan Song, Sipeng Zheng, and Qin Jin. In ICLR 2025. [code]
-
Modeling Fine-Grained Hand-Object Dynamics for Egocentric Video Representation Learning - Baoqi Pei, Yifei Huang, Jilan Xu, Guo Chen, Yuping He, Lijin Yang, Yali Wang, Weidi Xie, Yu Qiao, Fei Wu, and Limin Wang. In ICLR 2025. [code]
-
HourVideo: 1-Hour Video-Language Understanding - Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. In NeurIPS 2024. [project page] [code]
-
HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language Model - Khoa Vo, Thinh Phan, Kashu Yamazaki, Minh Tran, and Ngan Le. In NeurIPS 2024. [project page] [code]
-
AMEGO: Active Memory from long EGOcentric videos - Gabriele Goletto, Tushar Nagarajan, Giuseppe Averta, and Dima Damen. In ECCV 2024. [project page] [code]
-
Video ReCap: Recursive Captioning of Hour-Long Videos - Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, and Gedas Bertasius. In CVPR 2024. [project page]
-
Retrieval-Augmented Egocentric Video Captioning - Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie. In CVPR 2024. [project page]
-
Grounded Question-Answering in Long Egocentric Videos - Shangzhe Di and Weidi Xie. In CVPR 2024. [project page] [code]
-
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models - Sijie Cheng, Zhicheng Guo, Jingwen Wu, Kechen Fang, Peng Li, Huaping Liu, and Yang Liu. In CVPR 2024. [code]
-
EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation - Baoqi Pei, Guo Chen, Jilan Xu, Yuping He, Yicheng Liu, Kanghua Pan, et al. arXiv 2024. [code]
-
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding - Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. In NeurIPS 2023. [project page] [code]
Show papers (13)
-
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses - Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy, Umar Iqbal, and Juergen Gall. In CVPR 2026. [project page]
-
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos - Mengmeng Ge, Takashi Isobe, Xu Jia, Yanan Sun, Zetong Yang, Weinong Wang, Dong Zhou, Dong Li, Huchuan Lu, and Emad Barsoum. In CVPR 2026.
-
EgoX: Egocentric Video Generation from a Single Exocentric Video - Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park, Junha Hyung, and Jaegul Choo. In CVPR 2026. [project page] [code]
-
Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos - Yujin Ham, Junho Kim, Vivek Boominathan, and Guha Balakrishnan. In CVPR 2026.
-
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing - Runjia Li, Moayed Haji-Ali, Ashkan Mirzaei, Chaoyang Wang, Arpit Sahni, Ivan Skorokhodov, Aliaksandr Siarohin, Tomas Jakab, Junlin Han, Sergey Tulyakov, Philip Torr, and Willi Menapace. In CVPR 2026. [project page] [code]
-
PlayerOne: Egocentric World Simulator - Yuanpeng Tu, Hao Luo, Xi Chen, Xiang Bai, Fan Wang, and Hengshuang Zhao. In NeurIPS 2025. [project page] [code]
-
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation - Xiaofeng Wang, Kang Zhao, Feng Liu, Jiayu Wang, Guosheng Zhao, Xiaoyi Bao, Zheng Zhu, Yingya Zhang, and Xingang Wang. In NeurIPS 2025. [project page] [code]
-
Whole-Body Conditioned Egocentric Video Prediction - Yutong Bai, Danny Tran, Amir Bar, Yann LeCun, Trevor Darrell, and Jitendra Malik. In NeurIPS 2025. [project page]
-
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds - Lu Chen, Yizhou Wang, Shixiang Tang, Qianhong Ma, Tong He, Wanli Ouyang, Xiaowei Zhou, Hujun Bao, and Sida Peng. In ICCV 2025. [code]
-
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control - Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, et al. In CVPR 2025. [project page] [code]
-
Exocentric-to-Egocentric Video Generation - Jia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, and Mike Zheng Shou. In NeurIPS 2024. [code]
-
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning - Bolin Lai, Xiaoliang Dai, Lawrence Chen, Guan Pang, James M. Rehg, and Miao Liu. In ECCV 2024. [project page]
-
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos - Mi Luo, Zihui Xue, Alex Dimakis, and Kristen Grauman. In ECCV 2024. [project page]
Show papers (8)
-
FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos - Alexandros Delitzas, Chenyangguang Zhang, Alexey Gavryushin, Tommaso Di Mario, Boyang Sun, Rishabh Dabral, Leonidas Guibas, Christian Theobalt, Marc Pollefeys, Francis Engelmann, and Daniel Barath. In CVPR 2026. [project page]
-
Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset - Zirui Wang, Wenjing Bian, Xinghui Li, Yifu Tao, Jianeng Wang, Maurice Fallon, and Victor Adrian Prisacariu. In NeurIPS 2025. [project page]
-
Pandora: Articulated 3D Scene Graphs from Egocentric Vision - Alan Yu, Yun Chang, Christopher Xie, and Luca Carlone. In BMVC 2025.
-
Benchmarking Egocentric Visual-Inertial SLAM at City Scale - Anusha Krishnan, Shaohui Liu, Paul-Edouard Sarlin, Oscar Gentilhomme, David Caruso, Maurizio Monge, Richard Newcombe, Jakob Engel, and Marc Pollefeys. In ICCV 2025. [project page] [code]
-
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos - Chengbo Yuan, Geng Chen, Li Yi, and Yang Gao. In ICCV 2025. [project page]
-
DIV-FF: Dynamic Image-Video Feature Fields for Environment Understanding in Egocentric Videos - Lorenzo Mur-Labadia, Jose J. Guerrero, and Ruben Martinez-Cantin. In CVPR 2025. [code]
-
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos - Vadim Tschernezki, Diane Larlus, Iro Laina, and Andrea Vedaldi. In CVPR 2025.
-
EgoLifter: Open-world 3D Segmentation for Egocentric Perception - Qiao Gu, Zhaoyang Lv, Duncan Frost, Simon Green, Julian Straub, Chris Sweeney, et al. In ECCV 2024. [project page]
Show papers (6)
-
LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks - Hengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen, Qihang Cao, Xianfeng Wang, Yucheng Zhu, Xiongkuo Min, Wei Sun, Dandan Zhu, and Guangtao Zhai. In CVPR 2026.
-
Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents - Vijay Veerabadran, Fanyi Xiao, Nitin Kamra, Pedro Matias, Joy Chen, Caley Drooff, et al. In NeurIPS 2025.
-
EgoBlind: Towards Egocentric Visual Assistance for the Blind - Junbin Xiao, Nanxin Huang, Hao Qiu, Zhulin Tao, Xun Yang, Richang Hong, Meng Wang, and Angela Yao. In NeurIPS 2025.
-
LookOut: Real-World Humanoid Egocentric Navigation - Boxiao Pan, Adam W. Harley, Francis Engelmann, C. Karen Liu, and Leonidas J. Guibas. In ICCV 2025. [project page]
-
Vid2Coach: Transforming How-To Videos into Task Assistants - Mina Huh, Zihui Xue, Ujjaini Das, Kumar Ashutosh, Kristen Grauman, and Amy Pavel. In UIST 2025. [project page]
-
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset - Sagar M. Waghmare, Kimberly Wilber, Dave Hawkey, Xuan Yang, Matthew Wilson, Stephanie Debats, et al. In WACV 2025. [project page] [code]
Show papers (42)
-
EgoPet: Egomotion and Interaction Data from an Animal's Perspective - Amir Bar, Arya Bakhtiar, Danny Tran, Antonio Loquercio, Jathushan Rajasegaran, Yann LeCun, Amir Globerson, and Trevor Darrell. In ECCV 2024. [project page]
-
egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks - Matthias Jammot, Björn Braun, Paul Streli, Rafael Wampfler, and Christian Holz. In NeurIPS 2025. [project page] [code]
-
egoPPG: Heart Rate Estimation from Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision Tasks - Björn Braun, Rayan Armani, Manuel Meier, Max Moebus, and Christian Holz. In ICCV 2025. [project page] [code]
-
Toward Robust Audio-Visual Synchronization Detection in Egocentric Video with Sparse Synchronization Events - Jordan Voas, Wei-Cheng Tseng, Benoit Vallade, Alex Mackin, David Higham, and David Harwath. In BMVC 2025.
-
SkillSight: Efficient First-Person Skill Assessment with Gaze - Chi Hsuan Wu, Kumar Ashutosh, and Kristen Grauman. In CVPR 2026.
-
PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric Videos - Xun Jiang, Zhiyi Huang, Xing Xu, Jingkuan Song, Fumin Shen, and Heng Tao Shen. In CVPR 2025.
-
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos - Aashish Rai and Srinath Sridhar. In WACV 2025. [project page]
-
E³: Exploring Embodied Emotion Through A Large-Scale Egocentric Video Dataset - Wang Lin, Yueying Feng, Wenkang Han, Tao Jin, Zhou Zhao, Fei Wu, Chang Yao, and Jingyuan Chen. In NeurIPS 2024.
-
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval - Thomas Hummel, Shyamgopal Karthik, Mariana-Iuliana Georgescu, and Zeynep Akata. In ECCV 2024. [code]
-
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos - Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, and Kristen Grauman. In ECCV 2024. [project page] [code]
-
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition - Mingfang Zhang, Yifei Huang, Ruicong Liu, and Yoichi Sato. In ECCV 2024. [code]
-
EgoGen: An Egocentric Synthetic Data Generator - Gen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu, Mihai Dusmanu, Yan Zhang, Marc Pollefeys, and Siyu Tang. In CVPR 2024. [project page]
-
PREGO: online mistake detection in PRocedural EGOcentric videos - Alessandro Flaborea, Guido Maria D'Amely di Melendugno, Leonardo Plini, Luca Scofano, Edoardo De Matteis, Antonino Furnari, Giovanni Maria Farinella, and Fabio Galasso. In CVPR 2024.
-
Action Scene Graphs for Long-Form Understanding of Egocentric Videos - Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi, and Giovanni Maria Farinella. In CVPR 2024.
-
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos - Sagnik Majumder, Ziad Al-Halah, and Kristen Grauman. In CVPR 2024. [project page]
-
Trans4Map: Revisiting Holistic Bird's-Eye-View Mapping from Egocentric Images to Allocentric Semantics with Vision Transformers - Chang Chen, Jiaming Zhang, Kailun Yang, Kunyu Peng, and Rainer Stiefelhagen. In WACV 2023.
-
Multi-label Affordance Mapping from Egocentric Vision - Lorenzo Mur-Labadia, Jose J. Guerrero, and Ruben Martinez-Cantin. In ICCV 2023.
-
COPILOT: Human-Environment Collision Prediction and Localization from Egocentric Videos - Boxiao Pan, Bokui Shen, Davis Rempe, Despoina Paschalidou, Kaichun Mo, Yanchao Yang, and Leonidas J. Guibas. In ICCV 2023. [project page]
-
Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition - Qitong Wang, Long Zhao, Liangzhe Yuan, Ting Liu, and Xi Peng. In ICCV 2023. [code]
-
Tracking Multiple Deformable Objects in Egocentric Videos - Mingzhen Huang, Xiaoxing Li, Jun Hu, Honghong Peng, Siwei Lyu. In CVPR 2023.
-
Egocentric Audio-Visual Object Localization - Chao Huang · Yapeng Tian · Anurag Kumar · Chenliang Xu. In CVPR 2023. [project page]
-
Balanced Spherical Grid for Egocentric View Synthesis - Changwoon Choi · Sang Min Kim · Young Min Kim. In CVPR 2023. [code]
-
Ego-Body Pose Estimation via Ego-Head Pose Estimation - Jiaman Li · Karen Liu · Jiajun Wu. In CVPR 2023.
-
Egocentric Video Task Translation - Zihui Xue · Yale Song · Kristen Grauman · Lorenzo Torresani. In CVPR 2023.
-
Egocentric Auditory Attention Localization in Conversations - Fiona Ryan · Hao Jiang · Abhinav Shukla · James Rehg · Vamsi Krishna Ithapu. In CVPR 2023. [project page]
-
Where is my Wallet? Modeling Object Proposal Sets for Egocentric Visual Query Localization - Mengmeng Xu · Yanghao Li · Cheng-Yang Fu · Bernard Ghanem · Tao Xiang · Juan-Manuel Perez-Rua. In CVPR 2023. [project page]
-
Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations - Sagnik Majumder · Hao Jiang · Pierre Moulon · Ethan Henderson · Paul Calamia · Kristen Grauman · Vamsi Krishna Ithapu. In CVPR 2023.
-
EgoTaskQA: Understanding Human Tasks in Egocentric Videos - Baoxiong Jia, Ting Lei, Song-Chun Zhu, Siyuan Huang. In NeurIPS 2022. [projet page] [code]
-
Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Reality - Amin Jourabloo, Fernando De la Torre, Jason Saragih, Shih-En Wei, Stephen Lombardi, Te-Li Wang, Danielle Belko, Autumn Trimble, Hernan Badino. In CVPR 2022.
-
Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos - Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong Wang. In CVPR 2022. [project page] [video] [slides]
-
Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization - Hao Jiang, Calvin Murdock, Vamsi Krishna Ithapu. In CVPR 2022.
-
Egocentric Scene Understanding via Multimodal Spatial Rectifier - Tien Do, Khiem Vuong, Hyun Soo Park. In CVPR 2022.
-
Egocentric Prediction of Action Target in 3D - Yiming Li, Ziang Cao, Andrew Liang, Benjamin Liang, Luoyao Chen, Hang Zhao, Chen Feng. In CVPR 2022.
-
Slow-Fast Auditory Streams for Audio Recognition - Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. ICASSP 2021. [project page] [code]
-
Ego-Exo: Transferring Visual Representations From Third-Person to First-Person Videos - Yanghao Li, Tushar Nagarajan, Bo Xiong, Kristen Grauman. In CVPR 2021. [code]
-
EGO-SLAM: A Robust Monocular SLAM for Egocentric Videos - Suvam Patra, Kartikeya Gupta, Faran Ahmad, Chetan Arora, and Subhashis Banerjee. In WACV 2019. [code]
-
Egocentric Basketball Motion Planning from a Single First-Person Image - Gedas Bertasius, Aaron Chan, and Jianbo Shi. In CVPR 2018. [demo]
-
Jointly Learning Energy Expenditures and Activities using Egocentric Multimodal Signals - Katsuyuki Nakamura, Serena Yeung, Alexandre Alahi, and Li Fei-Fei. In CVPR 2017.
-
Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual Data - Jing Wang, Yu Cheng, and Rogerio Schmidt Feris. In CVPR 2016. [demo]
-
Compact CNN for Indexing Egocentric Videos - Yair Poleg, Ariel Ephrat, Shmuel Peleg, and Chetan Arora. In WACV 2016.
-
Detecting engagement in egocentric video - Yu-Chuan Su and Kristen Grauman. In ECCV 2016.
-
EgoSampling: Fast-Forward and Stereo for Egocentric Videos - Yair Poleg, Tavi Halperin, Chetan Arora, and Shmuel Peleg. In CVPR 2015.
Clustered according to the conferences.
Show papers (147)
-
ForeHOI: Feed-forward 3D Object Reconstruction from Daily Hand-Object Interaction Videos - Yuantao Chen, Jiahao Chang, Chongjie Ye, Chaoran Zhang, Zhaojie Fang, Chenghong Li, and Xiaoguang Han. In CVPR 2026. [project page] [code]
-
EgoFlow: Gradient-Guided Flow Matching for Egocentric 6DoF Object Motion Generation - Abhishek Saroha, Huajian Zeng, Xingxing Zuo, Daniel Cremers, and Xi Wang. In CVPR 2026. [project page] [code]
-
EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive Hierarchy - Jinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan, Yifan Yu, Dong Wang, Honglei Yan, Liang Yue, Shaofei Wang, Yixin Chen, Siyuan Huang, and Miao Liu. In CVPR 2026.
-
Ego-Grounding for Personalized Question-Answering in Egocentric Videos - Junbin Xiao, Shenglang Zhang, Pengxiang Zhu, and Angela Yao. In CVPR 2026. [code]
-
EgoAVU: Egocentric Audio-Visual Understanding - Ashish Seth, Xinhao Mei, Changsheng Zhao, Varun Nagaraja, Ernie Chang, Gregory P. Meyer, Gael Le Lan, Yunyang Xiong, Vikas Chandra, Yangyang Shi, Dinesh Manocha, and Zhipeng Cai. In CVPR 2026. [project page] [code]
-
EgoSound: Benchmarking Sound Understanding in Egocentric Videos - Bingwen Zhu, Yuqian Fu, Qiaole Dong, Guolei Sun, Tianwen Qian, Yuzheng Wu, Danda Pani Paudel, Xiangyang Xue, and Yanwei Fu. In CVPR 2026. [project page] [dataset]
-
Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding - Arsha Nagrani, Jasper Uijlings, Shyamal Buch, Tobias Weyand, Sudheendra Vijayanarasimhan, Bo Hu, Ramin Mehran, David A Ross, and Cordelia Schmid. In CVPR 2026. [dataset]
-
Do You See What I Am Pointing At? Gesture-Based Egocentric Video Question Answering - Yura Choi, Roy Miles, Rolandos Alexandros Potamias, Ismail Elezi, Jiankang Deng, and Stefanos Zafeiriou. In CVPR 2026. [project page]
-
Ego-InBetween: Generating Object State Transitions in Ego-Centric Videos - Mengmeng Ge, Takashi Isobe, Xu Jia, Yanan Sun, Zetong Yang, Weinong Wang, Dong Zhou, Dong Li, Huchuan Lu, and Emad Barsoum. In CVPR 2026.
-
EgoX: Egocentric Video Generation from a Single Exocentric Video - Taewoong Kang, Kinam Kim, Dohyeon Kim, Minho Park, Junha Hyung, and Jaegul Choo. In CVPR 2026. [project page] [code]
-
Generating Humanless Environment Walkthroughs from Egocentric Walking Tour Videos - Yujin Ham, Junho Kim, Vivek Boominathan, and Guha Balakrishnan. In CVPR 2026.
-
EgoEdit: Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing - Runjia Li, Moayed Haji-Ali, Ashkan Mirzaei, Chaoyang Wang, Arpit Sahni, Ivan Skorokhodov, Aliaksandr Siarohin, Tomas Jakab, Junlin Han, Sergey Tulyakov, Philip Torr, and Willi Menapace. In CVPR 2026. [project page] [code]
-
Mistake Attribution: Fine-Grained Mistake Understanding in Egocentric Videos - Yayuan Li, Aadit Jain, Filippos Bellos, and Jason J. Corso. In CVPR 2026. [project page]
-
Egocentric Visibility-Aware Human Pose Estimation - Peng Dai, Yu Zhang, Yiqiang Feng, Zhen Fan, and Yang Zhang. In CVPR 2026.
-
EgoPoseFormer v2: Accurate Egocentric Human Motion Estimation for AR/VR - Zhenyu Li, Sai Kumar Dwivedi, Filip Maric, Carlos Chacon, Nadine Bertsch, Filippo Arcadu, et al. In CVPR 2026. [project page]
-
Towards Stable Self-Supervised Object Representations in Unconstrained Egocentric Video - Yuting Tan, Xilong Cheng, Yunxiao Qin, Zhengnan Li, and Jingjing Zhang. In CVPR 2026.
-
EgoXtreme: A Dataset for Robust Object Pose Estimation in Egocentric Views under Extreme Conditions - Taegyoon Yoon, Yegyu Han, Seojin Ji, Jaewoo Park, Sojeong Kim, Taein Kwon, and Hyung-Sin Kim. In CVPR 2026. [project page] [code]
-
Test-time Ego-Exo-centric Adaptation for Action Anticipation via Multi-Label Prototype Growing and Dual-Clue Consistency - Zhaofeng Shi, Heqian Qiu, Lanxiao Wang, Qingbo Wu, Fanman Meng, Lili Pan, and Hongliang Li. In CVPR 2026. [code]
-
Forecasting 3D Scanpaths in Egocentric Video - Fiona Ryan, Ishwarya Ananthabhotla, Yijun Qian, Judy Hoffman, James M. Rehg, Vamsi Krishna Ithapu, and Calvin Murdock. In CVPR 2026.
-
SAVA-X: Ego-to-Exo Imitation Error Detection via Scene-Adaptive View Alignment and Bidirectional Cross View Fusion - Xiang Li, Heqian Qiu, Lanxiao Wang, Benliu Qiu, Fanman Meng, Linfeng Xu, and Hongliang Li. In CVPR 2026. [code]
-
Seeing Conversations: Communication Context Identification in Egocentric Video - Tobias Dorszewski and Jens Hjortkjær. In CVPR 2026.
-
LifeEval: A Multimodal Benchmark for Assistive AI in Egocentric Daily Life Tasks - Hengjian Gao, Kaiwei Zhang, Shibo Wang, Mingjie Chen, Qihang Cao, Xianfeng Wang, Yucheng Zhu, Xiongkuo Min, Wei Sun, Dandan Zhu, and Guangtao Zhai. In CVPR 2026.
-
SkillSight: Efficient First-Person Skill Assessment with Gaze - Chi Hsuan Wu, Kumar Ashutosh, and Kristen Grauman. In CVPR 2026.
-
EgoControl: Controllable Egocentric Video Generation via 3D Full-Body Poses - Enrico Pallotta, Sina Mokhtarzadeh Azar, Lars Doorenbos, Serdar Ozsoy, Umar Iqbal, and Juergen Gall. In CVPR 2026. [project page]
-
E-3DPSM: A State Machine for Event-Based Egocentric 3D Human Pose Estimation - Mayur Deshmukh, Hiroyasu Akada, Helge Rhodin, Christian Theobalt, and Vladislav Golyanik. In CVPR 2026. [project page]
-
Beyond Caption-Based Queries for Video Moment Retrieval - David Pujol-Perich, Albert Clapés, Dima Damen, Sergio Escalera, and Michael Wray. In CVPR 2026.
-
Ego2Web: A Web Agent Benchmark Grounded in Egocentric Videos - Shoubin Yu, Lei Shu, Antoine Yang, Yao Fu, Srinivas Sunkara, Maria Wang, Jindong Chen, Mohit Bansal, and Boqing Gong. In CVPR 2026. [project page] [code]
-
HanDyVQA: A Video QA Benchmark for Fine-Grained Hand-Object Interaction Dynamics - Masatoshi Tateno, Gido Kato, Hirokatsu Kataoka, Yoichi Sato, and Takuma Yagi. In CVPR 2026. [project page]
-
FunRec: Reconstructing Functional 3D Scenes from Egocentric Interaction Videos - Alexandros Delitzas, Chenyangguang Zhang, Alexey Gavryushin, Tommaso Di Mario, Boyang Sun, Rishabh Dabral, Leonidas Guibas, Christian Theobalt, Marc Pollefeys, Francis Engelmann, and Daniel Barath. In CVPR 2026. [project page]
-
Object-Shot Enhanced Grounding Network for Egocentric Video - Yisen Feng, Haoyu Zhang, Meng Liu, Weili Guan, and Liqiang Nie. In CVPR 2025. [code]
-
ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition Benchmark - Ronghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin, Boqiang Zhang, Long Li, Liuyi Wang, Qinyang Zeng, Xin Li, and Lidong Bing. In CVPR 2025. [code]
-
ANNEXE: Unified Analyzing, Answering, and Pixel Grounding for Egocentric Interaction - Yuejiao Su, Yi Wang, Qiongyang Hu, Chuang Yang, and Lap-Pui Chau. In CVPR 2025. [project page]
-
Understanding Multi-Task Activities from Single-Task Videos - Yuhan Shen and Ehsan Elhamifar. In CVPR 2025.
-
Ego4o: Egocentric Human Motion Capture and Understanding from Multi-Modal Input - Jian Wang, Rishabh Dabral, Diogo Luvizon, Zhe Cao, Lingjie Liu, Thabo Beeler, and Christian Theobalt. In CVPR 2025. [project page]
-
Sound Bridge: Associating Egocentric and Exocentric Videos via Audio Cues - Sihong Huang, Jiaxin Wu, Xiaoyong Wei, Yi Cai, Dongmei Jiang, and Yaowei Wang. In CVPR 2025. [code]
-
PHGC: Procedural Heterogeneous Graph Completion for Natural Language Task Verification in Egocentric Videos - Xun Jiang, Zhiyi Huang, Xing Xu, Jingkuan Song, Fumin Shen, and Heng Tao Shen. In CVPR 2025.
-
Video ReCap: Recursive Captioning of Hour-Long Videos - Md Mohaiminul Islam, Ngan Ho, Xitong Yang, Tushar Nagarajan, Lorenzo Torresani, and Gedas Bertasius. In CVPR 2024. [project page]
-
Retrieval-Augmented Egocentric Video Captioning - Jilan Xu, Yifei Huang, Junlin Hou, Guo Chen, Yuejie Zhang, Rui Feng, and Weidi Xie. In CVPR 2024. [project page]
-
Grounded Question-Answering in Long Egocentric Videos - Shangzhe Di and Weidi Xie. In CVPR 2024. [project page] [code]
-
Learning to Segment Referred Objects from Narrated Egocentric Videos - Yuhan Shen, Huiyu Wang, Xitong Yang, Matt Feiszli, Ehsan Elhamifar, Lorenzo Torresani, and Effrosyni Mavroudi. In CVPR 2024.
-
Error Detection in Egocentric Procedural Task Videos - Shih-Po Lee, Zijia Lu, Zekun Zhang, Minh Hoai, and Ehsan Elhamifar. In CVPR 2024. [project page] [code]
-
Progress-Aware Online Action Segmentation for Egocentric Procedural Task Videos - Yuhan Shen and Ehsan Elhamifar. In CVPR 2024. [code]
-
TIM: A Time Interval Machine for Audio-Visual Action Recognition - Jacob Chalk, Jaesung Huh, Evangelos Kazakos, Andrew Zisserman, and Dima Damen. In CVPR 2024. [project page] [code]
-
Real-Time Simulated Avatar from Head-Mounted Sensors - Zhengyi Luo, Jinkun Cao, Rawal Khirodkar, Alexander Winkler, Jing Huang, Kris Kitani, and Weipeng Xu. In CVPR 2024. [project page]
-
Mocap Everyone Everywhere: Lightweight Motion Capture With Smartwatches and a Head-Mounted Camera - Jiye Lee and Hanbyul Joo. In CVPR 2024. [project page]
-
Fusing Personal and Environmental Cues for Identification and Segmentation of First-Person Camera Wearers in Third-Person Views - Ziwei Zhao, Yuchen Wang, Chuhua Wang, and David Crandall. In CVPR 2024. [code]
-
Learning Spatial Features from Audio-Visual Correspondence in Egocentric Videos - Sagnik Majumder, Ziad Al-Halah, and Kristen Grauman. In CVPR 2024. [project page]
-
REWIND: Real-Time Egocentric Whole-Body Motion Diffusion with Exemplar-Based Identity Conditioning - Jihyun Lee, Weipeng Xu, Alexander Richard, Shih-En Wei, Shunsuke Saito, Shaojie Bai, Te-Li Wang, Minhyuk Sung, Tae-Kyun Kim, and Jason Saragih. In CVPR 2025. [project page]
-
FRAME: Floor-aligned Representation for Avatar Motion from Egocentric Video - Andrea Boscolo Camiletto, Jian Wang, Eduardo Alvarado, Rishabh Dabral, Thabo Beeler, Marc Habermann, and Christian Theobalt. In CVPR 2025. [project page] [code]
-
EgoLM: Multi-Modal Language Model of Egocentric Motions - Fangzhou Hong, Vladimir Guzov, Hyo Jin Kim, Yuting Ye, Richard Newcombe, Ziwei Liu, and Lingni Ma. In CVPR 2025. [project page]
-
Estimating Body and Hand Motion in an Ego-sensed World - Brent Yi, Vickie Ye, Maya Zheng, Yunqi Li, Lea Müller, Georgios Pavlakos, Yi Ma, Jitendra Malik, and Angjoo Kanazawa. In CVPR 2025. [project page]
-
Dyn-HaMR: Recovering 4D Interacting Hand Motion from a Dynamic Camera - Zhengdi Yu, Stefanos Zafeiriou, and Tolga Birdal. In CVPR 2025. [project page]
-
EgoPressure: A Dataset for Hand Pressure and Pose Estimation in Egocentric Vision - Yiming Zhao, Taein Kwon, Paul Streli, Marc Pollefeys, and Christian Holz. In CVPR 2025. [project page]
-
HaWoR: World-Space Hand Motion Reconstruction from Egocentric Videos - Jinglei Zhang, Jiankang Deng, Chao Ma, and Rolandos Alexandros Potamias. In CVPR 2025. [project page]
-
HOT3D: Hand and Object Tracking in 3D from Egocentric Multi-View Videos - Prithviraj Banerjee, Sindi Shkodrani, Pierre Moulon, Shreyas Hampali, Shangchen Han, Fan Zhang, et al. In CVPR 2025. [project page]
-
ParaHome: Parameterizing Everyday Home Activities Towards 3D Generative Modeling of Human-Object Interactions - Jeonghwan Kim, Jisoo Kim, Jeonghyeon Na, and Hanbyul Joo. In CVPR 2025. [code]
-
Generating 6DoF Object Manipulation Trajectories from Action Description in Egocentric Vision - Tomoya Yoshida, Shuhei Kurita, Taichi Nishimura, and Shinsuke Mori. In CVPR 2025. [project page] [code]
-
FIction: 4D Future Interaction Prediction from Video - Kumar Ashutosh, Georgios Pavlakos, and Kristen Grauman. In CVPR 2025. [code]
-
Gazing Into Missteps: Leveraging Eye-Gaze for Unsupervised Mistake Detection in Egocentric Videos of Skilled Human Activities - Michele Mazzamuto, Antonino Furnari, Yoichi Sato, and Giovanni Maria Farinella. In CVPR 2025.
-
HD-EPIC: A Highly-Detailed Egocentric Video Dataset - Toby Perrett, Ahmad Darkhalil, Saptarshi Sinha, Omar Emara, Sam Pollard, Kranti Parida, et al. In CVPR 2025. [project page]
-
GEM: A Generalizable Ego-Vision Multimodal World Model for Fine-Grained Ego-Motion, Object Dynamics, and Scene Composition Control - Mariam Hassan, Sebastian Stapf, Ahmad Rahimi, et al. In CVPR 2025. [project page] [code]
-
DIV-FF: Dynamic Image-Video Feature Fields for Environment Understanding in Egocentric Videos - Lorenzo Mur-Labadia, Jose J. Guerrero, and Ruben Martinez-Cantin. In CVPR 2025. [code]
-
Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric Videos - Vadim Tschernezki, Diane Larlus, Iro Laina, and Andrea Vedaldi. In CVPR 2025.
-
EgoLife: Towards Egocentric Life Assistant - Jingkang Yang, Shuai Liu, Hongming Guo, et al. In CVPR 2025. [code]
-
EgoTextVQA: Towards Egocentric Scene-Text Aware Video Question Answering - Sheng Zhou, Junbin Xiao, Qingyun Li, Yicong Li, Xun Yang, Dan Guo, Meng Wang, Tat-Seng Chua, and Angela Yao. In CVPR 2025. [code]
-
Omnia de EgoTempo: Benchmarking Temporal Understanding of Multi-Modal LLMs in Egocentric Videos - Chiara Plizzari, Alessio Tonioni, Yongqin Xian, Achin Kulshrestha, and Federico Tombari. In CVPR 2025. [code]
-
Bootstrap Your Own Views: Masked Ego-Exo Modeling for Fine-grained View-invariant Video Representations - Jungin Park, Jiyoung Lee, and Kwanghoon Sohn. In CVPR 2025. [code]
-
Viewpoint Rosetta Stone: Unlocking Unpaired Ego-Exo Videos for View-invariant Representation Learning - Mi Luo, Zihui Xue, Alex Dimakis, and Kristen Grauman. In CVPR 2025. [project page]
-
Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos - Sagnik Majumder, Tushar Nagarajan, Ziad Al-Halah, Reina Pradhan, and Kristen Grauman. In CVPR 2025. [project page]
-
EventEgo3D: 3D Human Motion Capture from Egocentric Event Streams - Christen Millerdurai, Hiroyasu Akada, Jian Wang, Diogo Luvizon, Christian Theobalt, and Vladislav Golyanik. In CVPR 2024. [project page]
-
3D Human Pose Perception from Egocentric Stereo Videos - Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. In CVPR 2024. [code]
-
Egocentric Whole-Body Motion Capture with FisheyeViT and Diffusion-Based Motion Refinement - Jian Wang, Zhe Cao, Diogo Luvizon, Lingjie Liu, Kripasindhu Sarkar, Danhang Tang, Thabo Beeler, and Christian Theobalt. In CVPR 2024. [code]
-
Attention-Propagation Network for Egocentric Heatmap to 3D Pose Lifting - Taeho Kang and Youngki Lee. In CVPR 2024. [code]
-
Single-to-Dual-View Adaptation for Egocentric 3D Hand Pose Estimation - Ruicong Liu, Takehiko Ohkawa, Mingfang Zhang, and Yoichi Sato. In CVPR 2024. [code]
-
SoundingActions: Learning How Actions Sound from Narrated Egocentric Videos - Changan Chen, Kumar Ashutosh, Rohit Girdhar, David Harwath, and Kristen Grauman. In CVPR 2024. [project page]
-
X-MIC: Cross-Modal Instance Conditioning for Egocentric Action Generalization - Anna Kukleva, Fadime Sener, Edoardo Remelli, Bugra Tekin, Eric Sauser, Bernt Schiele, and Shugao Ma. In CVPR 2024. [code]
-
Instance Tracking in 3D Scenes from Egocentric Videos - Yunhan Zhao, Haoyu Ma, Shu Kong, and Charless Fowlkes. In CVPR 2024. [code]
-
Summarize the Past to Predict the Future: Natural Language Descriptions of Context Boost Multimodal Object Interaction Anticipation - Razvan-George Pasca, Alexey Gavryushin, Muhammad Hamza, Yen-Ling Kuo, Kaichun Mo, Luc Van Gool, Otmar Hilliges, and Xi Wang. In CVPR 2024. [project page]
-
The Audio-Visual Conversational Graph: From an Egocentric-Exocentric Perspective - Wenqi Jia, Miao Liu, Hao Jiang, Ishwarya Ananthabhotla, James M. Rehg, Vamsi Krishna Ithapu, and Ruohan Gao. In CVPR 2024. [project page]
-
EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models - Sijie Cheng, Zhicheng Guo, Jingwen Wu, Kechen Fang, Peng Li, Huaping Liu, and Yang Liu. In CVPR 2024. [code]
-
EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World - Yifei Huang, Guo Chen, Jilan Xu, Mingfang Zhang, Lijin Yang, Baoqi Pei, et al. In CVPR 2024. [code]
-
EgoGen: An Egocentric Synthetic Data Generator - Gen Li, Kaifeng Zhao, Siwei Zhang, Xiaozhong Lyu, Mihai Dusmanu, Yan Zhang, Marc Pollefeys, and Siyu Tang. In CVPR 2024. [project page]
-
A Backpack Full of Skills: Egocentric Video Understanding with Diverse Task Perspectives - Simone Alberto Peirone, Francesca Pistilli, Antonio Alliegro, and Giuseppe Averta. In CVPR 2024.
-
Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives - Kristen Grauman, Andrew Westbury, Lorenzo Torresani, Kris Kitani, Jitendra Malik, Triantafyllos Afouras, Kumar Ashutosh, Vijay Baiyya, Siddhant Bansal, Bikram Boote, Eugene Byrne, Zach Chavis, Joya Chen, Feng Cheng, Fu-Jen Chu, Sean Crane, Avijit Dasgupta, Jing Dong, Maria Escobar, Cristhian Forigua, Abrham Gebreselasie, Sanjay Haresh, Jing Huang, Md Mohaiminul Islam, Suyog Jain, Rawal Khirodkar, Devansh Kukreja, Kevin J Liang, Jia-Wei Liu, Sagnik Majumder, Yongsen Mao, Miguel Martin, Effrosyni Mavroudi, Tushar Nagarajan, Francesco Ragusa, Santhosh Kumar Ramakrishnan, Luigi Seminara, Arjun Somayazulu, Yale Song, Shan Su, Zihui Xue, Edward Zhang, Jinxu Zhang, Angela Castillo, Changan Chen, Xinzhu Fu, Ryosuke Furuta, Cristina Gonzalez, Prince Gupta, Jiabo Hu, Yifei Huang, Yiming Huang, Weslie Khoo, Anush Kumar, Robert Kuo, Sach Lakhavani, Miao Liu, Mi Luo, Zhengyi Luo, Brighid Meredith, Austin Miller, Oluwatumininu Oguntola, Xiaqing Pan, Penny Peng, Shraman Pramanick, Merey Ramazanova, Fiona Ryan, Wei Shan, Kiran Somasundaram, Chenan Song, Audrey Southerland, Masatoshi Tateno, Huiyu Wang, Yuchen Wang, Takuma Yagi, Mingfei Yan, Xitong Yang, Zecheng Yu, Shengxin Cindy Zha, Chen Zhao, Ziwei Zhao, Zhifan Zhu, Jeff Zhuo, Pablo Arbelaez, Gedas Bertasius, David Crandall, Dima Damen, Jakob Engel, Giovanni Maria Farinella, Antonino Furnari, Bernard Ghanem, Judy Hoffman, C. V. Jawahar, Richard Newcombe, Hyun Soo Park, James M. Rehg, Yoichi Sato, Manolis Savva, Jianbo Shi, Mike Zheng Shou, and Michael Wray. In CVPR 2024. [project page]
-
PREGO: online mistake detection in PRocedural EGOcentric videos - Alessandro Flaborea, Guido Maria D'Amely di Melendugno, Leonardo Plini, Luca Scofano, Edoardo De Matteis, Antonino Furnari, Giovanni Maria Farinella, and Fabio Galasso. In CVPR 2024.
-
Action Scene Graphs for Long-Form Understanding of Egocentric Videos - Ivan Rodin, Antonino Furnari, Kyle Min, Subarna Tripathi, and Giovanni Maria Farinella. In CVPR 2024.
-
Therbligs In Action: Video Understanding through Motion Primitives - Eadom Dessalene, Michael Maynord, Cornelia Fermu ̈ller, Yiannis Aloimonos. In CVPR 2023. [project page]
-
Hierarchical Temporal Transformer for 3D Hand Pose Estimation and Action Recognition from Egocentric RGB Videos - Yilin Wen, Hao Pan, Lei Yang, Jia Pan, Taku Komura, Wenping Wang. In CVPR 2023. [Code]
-
MMG-Ego4D: Multimodal Generalization in Egocentric Action Recognition - Xinyu Gong, Sreyas Mohan, Naina Dhingra, Jean-Charles Bazin, YILEI LI, Zhangyang Wang, Rakesh Ranjan. In CVPR 2023.
-
AssemblyHands: Towards Egocentric Activity Understanding via 3D Hand Pose Estimation - Takehiko Ohkawa, Kun He, Fadime Sener, Tomas Hodan, LUAN TRAN, Cem Keskin. In CVPR 2023.
-
Scene-aware Egocentric 3D Human Pose Estimation - Jian Wang, Diogo Luvizon, Weipeng Xu, Lingjie Liu, Kripasindhu Sarkar, Christian Theobalt. In CVPR 2023.
-
Tracking Multiple Deformable Objects in Egocentric Videos - Mingzhen Huang, Xiaoxing Li, Jun Hu, Honghong Peng, Siwei Lyu. In CVPR 2023.
-
Egocentric Audio-Visual Object Localization - Chao Huang · Yapeng Tian · Anurag Kumar · Chenliang Xu. In CVPR 2023. [project page]
-
Balanced Spherical Grid for Egocentric View Synthesis - Changwoon Choi · Sang Min Kim · Young Min Kim. In CVPR 2023. [code]
-
Egocentric Video Task Translation - Zihui Xue · Yale Song · Kristen Grauman · Lorenzo Torresani. In CVPR 2023.
-
Egocentric Auditory Attention Localization in Conversations - Fiona Ryan · Hao Jiang · Abhinav Shukla · James Rehg · Vamsi Krishna Ithapu. In CVPR 2023. [project page]
-
Chat2Map: Efficient Scene Mapping from Multi-Ego Conversations - Sagnik Majumder · Hao Jiang · Pierre Moulon · Ethan Henderson · Paul Calamia · Kristen Grauman · Vamsi Krishna Ithapu. In CVPR 2023.
-
ARCTIC: A Dataset for Dexterous Bimanual Hand-Object Manipulation - Zicong Fan, Omid Taheri, Dimitrios Tzionas, Muhammed Kocabas, Manuel Kaufmann, Michael J. Black, Otmar Hilliges. In CVPR 2023. [code]
-
[Learning Video Representations from Large Language Models](https://arxiv.org/pdf/2212.04501.pdf; https://facebookresearch.github.io/LaViLa) - Yue Zhao, Ishan Misra, Philipp Krähenbühl, Rohit Girdhar. In CVPR 2023. [project page] [code] [demo]
-
Ego4D: Around the World in 3,000 Hours of Egocentric Video - Kristen Grauman, Andrew Westbury, Eugene Byrne, Zachary Chavis, Antonino Furnari, Rohit Girdhar, Jackson Hamburger, Hao Jiang, Miao Liu, Xingyu Liu, Miguel Martin, Tushar Nagarajan, Ilija Radosavovic, Santhosh Kumar Ramakrishnan, Fiona Ryan, Jayant Sharma, Michael Wray, Mengmeng Xu, Eric Zhongcong Xu, Chen Zhao, Siddhant Bansal, Dhruv Batra, Vincent Cartillier, Sean Crane, Tien Do, Morrie Doulaty, Akshay Erapalli, Christoph Feichtenhofer, Adriano Fragomeni, Qichen Fu, Christian Fuegen, Abrham Gebreselasie, Cristina Gonzalez, James Hillis, Xuhua Huang, Yifei Huang, Wenqi Jia, Weslie Khoo, Jachym Kolar, Satwik Kottur, Anurag Kumar, Federico Landini, Chao Li, Yanghao Li, Zhenqiang Li, Karttikeya Mangalam, Raghava Modhugu, Jonathan Munro, Tullie Murrell, Takumi Nishiyasu, Will Price, Paola Ruiz Puentes, Merey Ramazanova, Leda Sari, Kiran Somasundaram, Audrey Southerland, Yusuke Sugano, Ruijie Tao, Minh Vo, Yuchen Wang, Xindi Wu, Takuma Yagi, Yunyi Zhu, Pablo Arbelaez, David Crandall, Dima Damen, Giovanni Maria Farinella, Bernard Ghanem, Vamsi Krishna Ithapu, C.V. Jawahar, Hanbyul Joo, Kris Kitani, Haizhou Li, Richard Newcombe, Aude Oliva, Hyun Soo Park, James M. Rehg, Yoichi Sato, Jianbo Shi, Mike Zheng Shou, Antonio Torralba, Lorenzo Torresani, Mingfei Yan, and Jitendra Malik. In CVPR 2022. [Github] [project page] [video]
-
HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object Interaction - Yunze Liu, Yun Liu, Che Jiang, Kangbo Lyu, Weikang Wan, Hao Shen, Boqiang Liang, Zhoujie Fu, He Wang, Li Yi. In CVPR 2022. [project page] [video]
-
E2(GO)MOTION: Motion Augmented Event Stream for Egocentric Action Recognition - Chiara Plizzari, Mirco Planamente, Gabriele Goletto, Marco Cannici, Emanuele Gusso, Matteo Matteucci, Barbara Caputo. In CVPR 2022.
-
Estimating Egocentric 3D Human Pose in the Wild with External Weak Supervision - Jian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Diogo Luvizon, Christian Theobalt. In CVPR 2022. [project page]
-
Robust Egocentric Photo-realistic Facial Expression Transfer for Virtual Reality - Amin Jourabloo, Fernando De la Torre, Jason Saragih, Shih-En Wei, Stephen Lombardi, Te-Li Wang, Danielle Belko, Autumn Trimble, Hernan Badino. In CVPR 2022.
-
Joint Hand Motion and Interaction Hotspots Prediction from Egocentric Videos - Shaowei Liu, Subarna Tripathi, Somdeb Majumdar, Xiaolong Wang. In CVPR 2022. [project page] [video] [slides]
-
A Hybrid Egocentric Activity Anticipation Framework via Memory-Augmented Recurrent and One-shot Representation Forecasting - Tianshan Liu and Kin-Man Lam. In CVPR 2022.
-
Egocentric Deep Multi-Channel Audio-Visual Active Speaker Localization - Hao Jiang, Calvin Murdock, Vamsi Krishna Ithapu. In CVPR 2022.
-
Egocentric Scene Understanding via Multimodal Spatial Rectifier - Tien Do, Khiem Vuong, Hyun Soo Park. In CVPR 2022.
-
Egocentric Prediction of Action Target in 3D - Yiming Li, Ziang Cao, Andrew Liang, Benjamin Liang, Luoyao Chen, Hang Zhao, Chen Feng. In CVPR 2022.
-
Assembly101: A Large-Scale Multi-View Video Dataset for Understanding Procedural Activities - Fadime Sener, Dibyadip Chatterjee, Daniel Shelepov, Kun He, Dipika Singhania, Robert Wang, and Angela Yao. In CVPR 2022. [project page]
-
Ego-Exo: Transferring Visual Representations From Third-Person to First-Person Videos - Yanghao Li, Tushar Nagarajan, Bo Xiong, Kristen Grauman. In CVPR 2021. [code]
-
Multi-Modal Domain Adaptation for Fine-Grained Action Recognition - Jonathan Munro and Dima Damen. In CVPR 2020. [project page] [code]
-
Generalizing Hand Segmentation in Egocentric Videos with Uncertainty-Guided Model Adaptation - Minjie Cai, Feng Lu, and Yoichi Sato. In CVPR 2020. [code]
-
Multimodal Future Localization and Emergence Prediction for Objects in Egocentric View With a Reachability Prior - Osama Makansi, Ozgun Cicek, Kevin Buchicchio, and Thomas Brox. In CVPR 2020. [demo] [code] [project page]
-
EGO-TOPO: Environment Affordances from Egocentric Video - Tushar Nagarajan, Yanghao Li, Christoph Feichtenhofer, and Kristen Grauman. In CVPR 2020. [project page] [demo]
-
You2Me: Inferring Body Pose in Egocentric Video via First and Second Person Interactions - Evonne Ng, Donglai Xiang, Hanbyul Joo, and Kristen Grauman. In CVPR 2020. [demo] [project page] [dataset] [code]
-
LSTA: Long Short-Term Attention for Egocentric Action Recognition - Swathikiran Sudhakaran, Sergio Escalera, and Oswald Lanz. In CVPR 2019. [code]
-
H+O: Unified Egocentric Recognition of 3D Hand-Object Poses and Interactions - Bugra Tekin, Federica Bogo, and Marc Pollefeys. In CVPR 2019. [video]
-
Deep Dual Relation Modeling for Egocentric Interaction Recognition - Haoxin Li, Yijun Cai, and Wei-Shi Zheng. In CVPR 2019.
-
Egocentric Activity Recognition on a Budget - Rafael Possas, Sheila Pinto Caceres, and Fabio Ramos. In CVPR 2018. [demo]
-
From Lifestyle VLOGs to Everyday Interaction - David F. Fouhey, Weicheng Kuo, Alexei A. Efros, and Jitendra Malik. In CVPR 2018. [project page]
-
Actor and Observer: Joint Modeling of First and Third-Person Videos - Gunnar A. Sigurdsson, Abhinav Gupta, Cordelia Schmid, Ali Farhadi, and Karteek Alahari. In CVPR 2018. [code]
-
Analysis of Hand Segmentation in the Wild - Aisha Urooj Khan and Ali Borji. In CVPR 2018.
-
First-Person Hand Action Benchmark with RGB-D Videos and 3D Hand Pose Annotations - Guillermo Garcia-Hernando, Shanxin Yuan, Seungryul Baek, and Tae-Kyun Kim. In CVPR 2018. [project page] [code]
-
Egocentric Basketball Motion Planning from a Single First-Person Image - Gedas Bertasius, Aaron Chan, and Jianbo Shi. In CVPR 2018. [demo]
-
Query-focused video summarization: Dataset, evaluation, and a memory network based approach - Aidean Sharghi, Jacob S. Laurel and Boqing Gong. In CVPR 2017.
-
Deep future gaze: Gaze anticipation on egocentric videos using adversarial networks - Mengmi Zhang, Keng Teck Ma, Joo Hwee Lim, Qi Zhao, and Jiashi Feng. In CVPR 2017. [code]
-
Jointly Learning Energy Expenditures and Activities using Egocentric Multimodal Signals - Katsuyuki Nakamura, Serena Yeung, Alexandre Alahi, and Li Fei-Fei. In CVPR 2017.
-
Seeing Invisible Poses: Estimating 3D Body Pose from Egocentric Video - Hao Jiang and Kristen Grauman. In CVPR 2017.
-
First Person Action Recognition Using Deep Learned Descriptors - Suriya Singh, Chetan Arora, and C.V. Jawahar. In CVPR 2016. [project page] [code]
-
Going deeper into first-person activity recognition - Minghuang Ma, Haoqi Fan, and Kris M. Kitani. In CVPR 2016.
-
Egocentric Future Localization - Hyun Soo Park, Jyh-Jing Hwang, Yedong Niu, and Jianbo Shi. In CVPR 2016. [demo]
-
Recognizing Micro-Actions and Reactions from Paired Egocentric Videos - Ryo Yonetani, Kris M. Kitani, and Yoichi Sato. In CVPR 2016.
-
Walk and Learn: Facial Attribute Representation Learning from Egocentric Video and Contextual Data - Jing Wang, Yu Cheng, and Rogerio Schmidt Feris. In CVPR 2016. [demo]
-
Delving into egocentric actions - Yin Li, Zhefan Ye, and James M. Rehg. In CVPR 2015.
-
Pooled Motion Features for First-Person Videos - Michael S. Ryoo, Brandon Rothrock, and Larry H. Matthies. In CVPR 2015.
-
EgoSampling: Fast-Forward and Stereo for Egocentric Videos - Yair Poleg, Tavi Halperin, Chetan Arora, and Shmuel Peleg. In CVPR 2015.
-
Ego-Surfing First Person Videos - Ryo Yonetani, Kris M. Kitani, and Yoichi Sato. In CVPR 2015.
-
First-Person Pose Recognition using Egocentric Workspaces - Gregory Rogez, James S. Supancic, and Deva Ramanan. In CVPR 2015.
-
Temporal segmentation of egocentric videos -Yair Poleg, Chetan Arora, and Shmuel Peleg. In CVPR 2014.
-
First-Person Activity Recognition: What Are They Doing to Me? - M. S. Ryoo and Larry Matthies. In CVPR 2013.
-
Pixel-level hand detection in ego-centric videos - Cheng Li and Kris M. Kitani. In CVPR 2013. [video] [code]
-
Story-Driven Summarization for Egocentric Video - Zheng Lu and Kristen Grauman. In CVPR 2013 [project page]
-
Detecting activities of daily living in first-person camera views - Hamed Pirsiavash and Deva Ramanan. In CVPR 2012.
-
Discovering Important People and Objects for Egocentric Video Summarization - Yong Jae Lee, Joydeep Ghosh, and Kristen Grauman. In CVPR 2012. [project page]
-
Learning to recognize objects in egocentric activities - Alireza Fathi, Xiaofeng Ren, and James M. Rehg. In CVPR 2011.
-
Fast unsupervised ego-action learning for first-person sports videos - Kris M. Kitani, Takahiro Okabe, Yoichi Sato, and Akihiro Sugimoto. In CVPR 2011. [project page]
Show papers (42)
-
Towards in-the-wild Egocentric 3D Hand-Object Pose Estimation - Siddhant Bansal, Zhifan Zhu, Shashank Tripathi, Jiahe Zhao, Michael J. Black, and Dima Damen. In ECCV 2026. [project page] [code]
-
ActionVOS: Actions as Prompts for Video Object Segmentation - Liangyang Ouyang, Ruicong Liu, Yifei Huang, Ryosuke Furuta, and Yoichi Sato. In ECCV 2024. [code]
-
EgoBody3M: Egocentric Body Tracking on a VR Headset using a Diverse Dataset - Amy Zhao, Chengcheng Tang, Lezi Wang, Yijing Li, Mihika Dave, Lingling Tao, Christopher D. Twigg, and Robert Y. Wang. In ECCV 2024. [dataset]
-
Benchmarks and Challenges in Pose Estimation for Egocentric Hand Interactions with Objects - Zicong Fan, Takehiko Ohkawa, Linlin Yang, Nie Lin, Zhishan Zhou, Shihao Zhou, et al. In ECCV 2024.
-
Are Synthetic Data Useful for Egocentric Hand-Object Interaction Detection? - Rosario Leonardi, Antonino Furnari, Francesco Ragusa, and Giovanni Maria Farinella. In ECCV 2024. [project page] [code]
-
Spherical World-Locking for Audio-Visual Localization in Egocentric Videos - Heeseung Yun, Ruohan Gao, Ishwarya Ananthabhotla, Anurag Kumar, Jacob Donley, Chao Li, Gunhee Kim, Vamsi Krishna Ithapu, and Calvin Murdock. In ECCV 2024. [project page]
-
AFF-ttention! Affordances and Attention models for Short-Term Object Interaction Anticipation - Lorenzo Mur-Labadia, Ruben Martinez-Cantin, Jose J. Guerrero, Giovanni Maria Farinella, and Antonino Furnari. In ECCV 2024. [code]
-
PALM: Predicting Actions through Language Models - Sanghwan Kim, Daoji Huang, Yongqin Xian, Otmar Hilliges, Luc Van Gool, and Xi Wang. In ECCV 2024.
-
4Diff: 3D-Aware Diffusion Model for Third-to-First Viewpoint Translation - Feng Cheng, Mi Luo, Huiyu Wang, Alex Dimakis, Lorenzo Torresani, Gedas Bertasius, and Kristen Grauman. In ECCV 2024. [project page]
-
Synchronization is All You Need: Exocentric-to-Egocentric Transfer for Temporal Action Segmentation with Unlabeled Synchronized Video Pairs - Camillo Quattrocchi, Antonino Furnari, Daniele Di Mauro, Mario Valerio Giuffrida, and Giovanni Maria Farinella. In ECCV 2024. [code]
-
Ex2Eg-MAE: A Framework for Adaptation of Exocentric Video Masked Autoencoders for Egocentric Social Role Understanding - Minh Tran, Yelin Kim, Che-Chun Su, Cheng-Hao Kuo, Min Sun, and Mohammad Soleymani. In ECCV 2024.
-
EgoCVR: An Egocentric Benchmark for Fine-Grained Composed Video Retrieval - Thomas Hummel, Shyamgopal Karthik, Mariana-Iuliana Georgescu, and Zeynep Akata. In ECCV 2024. [code]
-
Action2Sound: Ambient-Aware Generation of Action Sounds from Egocentric Videos - Changan Chen, Puyuan Peng, Ami Baid, Zihui Xue, Wei-Ning Hsu, David Harwath, and Kristen Grauman. In ECCV 2024. [project page] [code]
-
Masked Video and Body-worn IMU Autoencoder for Egocentric Action Recognition - Mingfang Zhang, Yifei Huang, Ruicong Liu, and Yoichi Sato. In ECCV 2024. [code]
-
LEGO: Learning EGOcentric Action Frame Generation via Visual Instruction Tuning - Bolin Lai, Xiaoliang Dai, Lawrence Chen, Guan Pang, James M. Rehg, and Miao Liu. In ECCV 2024. [project page]
-
Put Myself in Your Shoes: Lifting the Egocentric Perspective from Exocentric Videos - Mi Luo, Zihui Xue, Alex Dimakis, and Kristen Grauman. In ECCV 2024. [project page]
-
EgoLifter: Open-world 3D Segmentation for Egocentric Perception - Qiao Gu, Zhaoyang Lv, Duncan Frost, Simon Green, Julian Straub, Chris Sweeney, et al. In ECCV 2024. [project page]
-
AMEGO: Active Memory from long EGOcentric videos - Gabriele Goletto, Tushar Nagarajan, Giuseppe Averta, and Dima Damen. In ECCV 2024. [project page] [code]
-
EgoExo-Fitness: Towards Egocentric and Exocentric Full-Body Action Understanding - Yuan-Ming Li, Wei-Jin Huang, An-Lan Wang, Ling-An Zeng, Jing-Ke Meng, and Wei-Shi Zheng. In ECCV 2024. [code]
-
Nymeria: A Massive Collection of Egocentric Multi-modal Human Motion in the Wild - Lingni Ma, Yuting Ye, Fangzhou Hong, Vladimir Guzov, Yifeng Jiang, et al. In ECCV 2024. [project page]
-
Listen to Look into the Future: Audio-Visual Egocentric Gaze Anticipation - Bolin Lai, Fiona Ryan, Wenqi Jia, Miao Liu, and James M. Rehg. In ECCV 2024. [project page]
-
On the Utility of 3D Hand Poses for Action Recognition - Md Salman Shamil, Dibyadip Chatterjee, Fadime Sener, Shugao Ma, and Angela Yao. In ECCV 2024. [project page]
-
EgoPoseFormer: A Simple Baseline for Stereo Egocentric 3D Human Pose Estimation - Chenhongyi Yang, Anastasia Tkach, Shreyas Hampali, Linguang Zhang, Elliot J. Crowley, and Cem Keskin. In ECCV 2024. [code]
-
EgoPoser: Robust Real-Time Egocentric Pose Estimation from Sparse and Intermittent Observations Everywhere - Jiaxi Jiang, Paul Streli, Manuel Meier, and Christian Holz. In ECCV 2024. [project page]
-
3D Hand Pose Estimation in Everyday Egocentric Images - Aditya Prakash, Ruisen Tu, Matthew Chang, and Saurabh Gupta. In ECCV 2024. [project page]
-
EgoPet: Egomotion and Interaction Data from an Animal's Perspective - Amir Bar, Arya Bakhtiar, Danny Tran, Antonio Loquercio, Jathushan Rajasegaran, Yann LeCun, Amir Globerson, and Trevor Darrell. In ECCV 2024. [project page]
-
My View is the Best View: Procedure Learning from Egocentric Videos - Siddhant Bansal, Chetan Arora, C.V. Jawahar. In ECCV 2022. [project page] [dataset] [code]
-
AssistQ: Affordance-centric Question-driven Task Completion for Egocentric Assistant - Benita Wong, Joya Chen, You Wu, Stan Weixian Lei, Dongxing Mao, Difei Gao, Mike Zheng Shou. In ECCV 2022. [project page] [code]
-
EgoBody: Human Body Shape and Motion of Interacting People from Head-Mounted Devices - Siwei Zhang, Qianli Ma, Yan Zhang, Zhiyin Qian, Taein Kwon, Marc Pollefeys, Federica Bogo, Siyu Tang. In ECCV 2022. [project page] [dataset] [code]
-
Generative Adversarial Network for Future Hand Segmentation from Egocentric Video - Wenqi Jia, Miao Liu, James M. Rehg. In ECCV 2022.
-
Fine-Grained Egocentric Hand-Object Segmentation: Dataset, Model, and Applications - Lingzhi Zhang, Shenghao Zhou, Simon Stent, Jianbo Shi. In ECCV 2022. [project page] [code] [dataset]
-
Egocentric Activity Recognition and Localization on a 3D Map - Miao Liu, Lingni Ma, Kiran Somasundaram, Yin Li, Kristen Grauman, James M. Rehg, Chao Li. In ECCV 2022.
-
SOS! Self-supervised Learning Over Sets Of Handled Objects In Egocentric Action Recognition - Victor Escorcia, Ricardo Guerrero, Xiatian Zhu, Brais Martinez. In ECCV 2022.
-
UnrealEgo: A New Dataset for Robust Egocentric 3D Human Motion Capture - Hiroyasu Akada, Jian Wang, Soshi Shimada, Masaki Takahashi, Christian Theobalt, Vladislav Golyanik. In ECCV 2022. [project page] [code] [dataset] [demo]
-
Forecasting Human-Object Interaction: Joint Prediction of Motor Attention and Actions in First Person Video - Miao Liu, Siyu Tang, Yin Li, and James M. Rehg. In ECCV 2020. [project page]
-
How Can I See My Future? FvTraj: Using First-person View for Pedestrian Trajectory Prediction - Huikun Bi, Ruisi Zhang, Tianlu Mao, Zhigang Deng, and Zhaoqi Wang. In ECCV 2020. [presentation video] [summary video]
-
Is Sharing of Egocentric Video Giving Away Your Biometric Signature? - Daksh Thapar, Chetan Arora, and Aditya Nigam. In ECCV 2020. [project page]
-
In the eye of beholder: Joint learning of gaze and actions in first person video - Yin Li, Miao Liu, and James M. Rehg. In ECCV 2018.
-
Predicting Gaze in Egocentric Video by Learning Task-dependent Attention Transition - Yifei Huang, Minjie Cai, Zhenqiang Li, and Yoichi Sato. In ECCV 2018 [code]
-
Detecting engagement in egocentric video - Yu-Chuan Su and Kristen Grauman. In ECCV 2016.
-
Detecting Snap Points in Egocentric Video with a Web Photo Prior - Bo Xiong and Kristen Grauman. In ECCV 2014. [project page] [code]
-
Learning to recognize daily actions using gaze - Alireza Fathi, Yin Li, and James M. Rehg. In ECCV 2012.
Show papers (56)
-
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding - Yue Fan, Xiaojian Ma, Rongpeng Su, Jun Guo, Rujie Wu, Xi Chen, and Qing Li. In ICCV 2025.
-
Visual Intention Grounding for Egocentric Assistants - Pengzhan Sun, Junbin Xiao, Tze Ho Elden Tse, Yicong Li, Arjun Akula, and Angela Yao. In ICCV 2025. [code]
-
EgoAgent: A Joint Predictive Agent Model in Egocentric Worlds - Lu Chen, Yizhou Wang, Shixiang Tang, Qianhong Ma, Tong He, Wanli Ouyang, Xiaowei Zhou, Hujun Bao, and Sida Peng. In ICCV 2025. [code]
-
EgoM2P: Egocentric Multimodal Multitask Pretraining - Gen Li, Yutong Chen, Yiqian Wu, Kaifeng Zhao, Marc Pollefeys, and Siyu Tang. In ICCV 2025. [project page]
-
O-MaMa: Learning Object Mask Matching between Egocentric and Exocentric Views - Lorenzo Mur-Labadia, Maria Santos-Villafranca, Jesus Bermudez-Cameo, Alejandro Perez-Yus, Ruben Martinez-Cantin, and Jose J. Guerrero. In ICCV 2025. [project page] [code]
-
Benchmarking Egocentric Visual-Inertial SLAM at City Scale - Anusha Krishnan, Shaohui Liu, Paul-Edouard Sarlin, Oscar Gentilhomme, David Caruso, Maurizio Monge, Richard Newcombe, Jakob Engel, and Marc Pollefeys. In ICCV 2025. [project page] [code]
-
Self-Supervised Monocular 4D Scene Reconstruction for Egocentric Videos - Chengbo Yuan, Geng Chen, Li Yi, and Yang Gao. In ICCV 2025. [project page]
-
egoPPG: Heart Rate Estimation from Eye-Tracking Cameras in Egocentric Systems to Benefit Downstream Vision Tasks - Björn Braun, Rayan Armani, Manuel Meier, Max Moebus, and Christian Holz. In ICCV 2025. [project page] [code]
-
PRVQL: Progressive Knowledge-guided Refinement for Robust Egocentric Visual Query Localization - Bing Fan, Yunhe Feng, Yapeng Tian, James Chenhao Liang, Yuewei Lin, Yan Huang, and Heng Fan. In ICCV 2025. [code]
-
Egocentric Action-aware Inertial Localization in Point Clouds with Vision-Language Guidance - Mingfang Zhang, Ryo Yonetani, Yifei Huang, Liangyang Ouyang, Ruicong Liu, and Yoichi Sato. In ICCV 2025.
-
Is Tracking Really More Challenging in First Person Egocentric Vision? - Matteo Dunnhofer, Zaira Manigrasso, and Christian Micheloni. In ICCV 2025. [project page] [code]
-
Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions - Liang Xu, Chengqun Yang, Zili Lin, Fei Xu, Yifan Liu, Congsheng Xu, et al. In ICCV 2025. [project page]
-
Learning Precise Affordances from Egocentric Videos for Robotic Manipulation - Gen Li, Nikolaos Tsagkas, Jifei Song, Ruaridh Mon-Williams, Sethu Vijayakumar, Kun Shao, and Laura Sevilla-Lara. In ICCV 2025. [project page]
-
ProbRes: Probabilistic Jump Diffusion for Open-World Egocentric Activity Recognition - Sanjoy Kundu, Shanmukha Vellamcheti, and Sathyanarayanan N. Aakur. In ICCV 2025.
-
Bring Your Rear Cameras for Egocentric 3D Human Pose Estimation - Hiroyasu Akada, Jian Wang, Vladislav Golyanik, and Christian Theobalt. In ICCV 2025. [project page]
-
Fish2Mesh Transformer: 3D Human Mesh Recovery from Egocentric Vision - Tianma Shen, Aditya Puranik, James Vong, Vrushabh Abhijit Deogirikar, Ryan Fell, Julianna Dietrich, Maria Kyrarini, Christopher Kitts, and David C. Jeong. In ICCV 2025. [project page]
-
EgoMusic-driven Human Dance Motion Estimation with Skeleton Mamba - Quang Nguyen, Nhat Le, Baoru Huang, Minh Nhat Vu, Chengcheng Tang, Van Nguyen, Ngan Le, Thieu Vo, and Anh Nguyen. In ICCV 2025.
-
UniEgoMotion: A Unified Model for Egocentric Motion Reconstruction, Forecasting, and Generation - Chaitanya Patel, Hiroki Nakamura, Yuta Kyuragi, Kazuki Kozuka, Juan Carlos Niebles, and Ehsan Adeli. In ICCV 2025. [project page] [code]
-
Head2Body: Body Pose Generation from Multi-sensory Head-mounted Inputs - Minh Tran, Hongda Mao, Qingshuang Chen, and Yelin Kim. In ICCV 2025.
-
ObjectRelator: Enabling Cross-View Object Relation Understanding Across Ego-Centric and Exo-Centric Perspectives - Yuqian Fu, Runze Wang, Bin Ren, Guolei Sun, Biao Gong, Yanwei Fu, Danda Pani Paudel, Xuanjing Huang, and Luc Van Gool. In ICCV 2025. [project page]
-
HiERO: Understanding the Hierarchy of Human Behavior Enhances Reasoning on Egocentric Videos - Simone Alberto Peirone, Francesca Pistilli, and Giuseppe Averta. In ICCV 2025. [project page] [code]
-
EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception - Sanjoy Chowdhury, Subrata Biswas, Sayan Nag, Tushar Nagarajan, Calvin Murdock, Ishwarya Ananthabhotla, Yijun Qian, Vamsi Krishna Ithapu, Dinesh Manocha, and Ruohan Gao. In ICCV 2025. [project page]
-
LookOut: Real-World Humanoid Egocentric Navigation - Boxiao Pan, Adam W. Harley, Francis Engelmann, C. Karen Liu, and Leonidas J. Guibas. In ICCV 2025. [project page]
-
Aria Digital Twin: A New Benchmark Dataset for Egocentric 3D Machine Perception - Xiaqing Pan, Nicholas Charron, Yongqian Yang, Scott Peters, Thomas Whelan, Chen Kong, Omkar Parkhi, Richard Newcombe, and Yuheng (Carl) Ren. In ICCV 2023. [project page]
-
EgoHumans: An Egocentric 3D Multi-Human Benchmark - Rawal Khirodkar, Aayush Bansal, Lingni Ma, Richard Newcombe, Minh Vo, and Kris Kitani. In ICCV 2023 (Oral). [code]
-
Self-Supervised Object Detection from Egocentric Videos - Peri Akiva, Jing Huang, Kevin J Liang, Rama Kovvuri, Xingyu Chen, Matt Feiszli, Kristin Dana, and Tal Hassner. In ICCV 2023.
-
Spectral Graphormer: Spectral Graph-Based Transformer for Egocentric Two-Hand Reconstruction using Multi-View Color Images - Tze Ho Elden Tse, Franziska Mueller, Zhengyang Shen, Danhang Tang, Thabo Beeler, Mingsong Dou, Yinda Zhang, Sasa Petrovic, Hyung Jin Chang, Jonathan Taylor, and Bardia Doosti. In ICCV 2023. [project page]
-
Uncertainty-aware State Space Transformer for Egocentric 3D Hand Trajectory Forecasting - Wentao Bao, Lele Chen, Libing Zeng, Zhong Li, Yi Xu, Junsong Yuan, and Yu Kong. In ICCV 2023. [project page] [code]
-
Probabilistic Human Mesh Recovery in 3D Scenes from Egocentric Views - Siwei Zhang, Qianli Ma, Yan Zhang, Sadegh Aliakbarian, Darren Cosker, and Siyu Tang. In ICCV 2023. [project page] [code]
-
HoloAssist: an Egocentric Human Interaction Dataset for Interactive AI Assistants in the Real World - Xin Wang, Taein Kwon, Mahdi Rad, Bowen Pan, Ishani Chakraborty, Sean Andrist, Dan Bohus, Ashley Feniello, Bugra Tekin, Felipe Vieira Frujeri, Neel Joshi, and Marc Pollefeys. In ICCV 2023. [project page]
-
EgoObjects: A Large-Scale Egocentric Dataset for Fine-Grained Object Understanding - Chenchen Zhu, Fanyi Xiao, Andres Alvarado, Yasmine Babaei, Jiabo Hu, Hichem El-Mohri, Sean Culatana, Roshan Sumbaly, and Zhicheng Yan. In ICCV 2023. [project page] [code]
-
EgoTV: Egocentric Task Verification from Natural Language Task Descriptions - Rishi Hazra, Brian Chen, Akshara Rai, Nitin Kamra, and Ruta Desai. In ICCV 2023. [project page] [code]
-
Ego-Only: Egocentric Action Detection without Exocentric Transferring - Huiyu Wang, Mitesh Kumar Singh, and Lorenzo Torresani. In ICCV 2023.
-
Multimodal Distillation for Egocentric Action Recognition - Gorjan Radevski, Dusan Grujicic, Matthew Blaschko, Marie-Francine Moens, and Tinne Tuytelaars. In ICCV 2023. [code]
-
Multi-label Affordance Mapping from Egocentric Vision - Lorenzo Mur-Labadia, Jose J. Guerrero, and Ruben Martinez-Cantin. In ICCV 2023.
-
COPILOT: Human-Environment Collision Prediction and Localization from Egocentric Videos - Boxiao Pan, Bokui Shen, Davis Rempe, Despoina Paschalidou, Kaichun Mo, Yanchao Yang, and Leonidas J. Guibas. In ICCV 2023. [project page]
-
Learning from Semantic Alignment between Unpaired Multiviews for Egocentric Video Recognition - Qitong Wang, Long Zhao, Liangzhe Yuan, Ting Liu, and Xi Peng. In ICCV 2023. [code]
-
EgoPCA: A New Framework for Egocentric Hand-Object Interaction Understanding - Yue Xu, Yong-Lu Li, Zhemin Huang, Michael Xu Liu, Cewu Lu, Yu-Wing Tai, and Chi-Keung Tang. In ICCV 2023. [project page]
-
EgoLoc: Revisiting 3D Object Localization from Egocentric Videos with Visual Queries - Jinjie Mai, Abdullah Hamdi, Silvio Giancola, Chen Zhao, and Bernard Ghanem. In ICCV 2023. [code]
-
What can a cook in Italy teach a mechanic in India? Action Recognition Generalisation Over Scenarios and Locations - Chiara Plizzari, Toby Perrett, Barbara Caputo, and Dima Damen. In ICCV 2023. [project page] [code]
-
EgoVLPv2: Egocentric Video-Language Pre-training with Fusion in the Backbone - Shraman Pramanick, Yale Song, Sayan Nag, Kevin Qinghong Lin, Hardik Shah, Mike Zheng Shou, Rama Chellappa, and Pengchuan Zhang. In ICCV 2023. [project page] [code]
-
Estimating Egocentric 3D Human Pose in Global Space - Jian Wang, Lingjie Liu, Weipeng Xu, Kripasindhu Sarkar, Christian Theobalt. In ICCV 2021. [project page]
-
Interactive Prototype Learning for Egocentric Action Recognition - Xiaohan Wang, Linchao Zhu, Heng Wang, and Yi Yang. In ICCV 2021.
-
EPIC-Fusion: Audio-Visual Temporal Binding for Egocentric Action Recognition - Evangelos Kazakos, Arsha Nagrani, Andrew Zisserman, and Dima Damen. In ICCV 2019. [code] [project page]
-
What Would You Expect? Anticipating Egocentric Actions with Rolling-Unrolling LSTMs and Modality Attention - Antonino Furnari and Giovanni Maria Farinella. In ICCV 2019 [code] [demo]
-
Ego-Pose Estimation and Forecasting as Real-Time PD Control - Ye Yuan and Kris Kitani. In ICCV 2019. [code] [project page] [demo]
-
xR-EgoPose: Egocentric 3D Human Pose From an HMD Camera - Denis Tome, Patrick Peluse, Lourdes Agapito, and Hernan Badino. In ICCV 2019. [demo] [dataset]
-
Jointly Recognizing Object Fluents and Tasks in Egocentric Videos - Yang Liu, Ping Wei, and Song-Chun Zhu. In ICCV 2017.
-
Egocentric Gesture Recognition Using Recurrent 3D Convolutional Neural Networks with Spatiotemporal Transformer Modules - Congqi Cao, Yifan Zhang, Yi Wu, Hanqing Lu, and Jian Cheng. In ICCV 2017.
-
First-Person Activity Forecasting with Online Inverse Reinforcement Learning - Nicholas Rhinehart and Kris M. Kitani. In ICCV 2017. [video]
-
Summarization and Classification of Wearable Camera Streams by Learning the Distributions over Deep Features of Out-of-Sample Image Sequences - Alessandro Perina, Sadegh Mohammadi, Nebojsa Jojic, and Vittorio Murino. In ICCV 2017.
-
Trespassing the Boundaries: Labeling Temporal Bounds for Object Interactions in Egocentric Video - Davide Moltisanti, Michael Wray, Walterio Mayol-Cuevas, and Dima Damen. In ICCV 2017.
-
Generating Notifications for Missing Actions: Don't forget to turn the lights off! - Bilge Soran, Ali Farhadi, and Linda Shapiro. In ICCV 2015.
-
Lending a hand: Detecting hands and recognizing activities in complex egocentric interactions - Sven Bambach, Stefan Lee, David J. Crandall, and Chen Yu. In ICCV 2015.
-
Learning to predict gaze in egocentric video - Yin Li, Alireza Fathi, and James M. Rehg. In ICCV 2013.
-
Context-based vision system for place and object recognition - Antonio Torralba, Kevin P. Murphy, William T. Freeman, Mark A. Rubin. In ICCV 2003. [project page]
Show papers (23)
-
ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos - Peiran Wu, Yunze Liu, Miao Liu, and Junxiao Shen. In WACV 2026.
-
Ego-EXTRA: video-language Egocentric Dataset for EXpert-TRAinee assistance - Francesco Ragusa, Michele Mazzamuto, Rosario Forte, Irene D'Ambra, James Fort, Jakob Engel, Antonino Furnari, and Giovanni Maria Farinella. In WACV 2026. [project page]
-
Towards Egocentric 3D Hand Pose Estimation in Unseen Domains - Wiktor Mucha, Michael Wray, and Martin Kampel. In WACV 2026.
-
RegionAligner: Bridging Ego-Exo Views for Object Correspondence via Unified Text-Visual Learning - Yuhao Su and Ehsan Elhamifar. In WACV 2026.
-
Online Episodic Memory Visual Query Localization with Egocentric Streaming Object Memory - Zaira Manigrasso, Matteo Dunnhofer, Antonino Furnari, Moritz Nottebaum, Antonio Finocchiaro, Davide Marana, Rosario Forte, Giovanni Maria Farinella, and Christian Micheloni. In WACV 2026.
-
Social EgoMesh Estimation - Luca Scofano, Alessio Sampieri, Edoardo De Matteis, Indro Spinelli, and Fabio Galasso. In WACV 2025. [code]
-
Ego-VPA: Egocentric Video Understanding with Parameter-Efficient Adaptation - Tz-Ying Wu, Kyle Min, Subarna Tripathi, and Nuno Vasconcelos. In WACV 2025. [code]
-
Exo2EgoDVC: Dense Video Captioning of Egocentric Procedural Activities Using Web Instructional Videos - Takehiko Ohkawa, Takuma Yagi, Taichi Nishimura, Ryosuke Furuta, Atsushi Hashimoto, Yoshitaka Ushiku, and Yoichi Sato. In WACV 2025. [code]
-
EgoSonics: Generating Synchronized Audio for Silent Egocentric Videos - Aashish Rai and Srinath Sridhar. In WACV 2025. [project page]
-
EgoCast: Forecasting Egocentric Human Pose in the Wild - Maria Escobar, Juanita Puentes, Cristhian Forigua, Jordi Pont-Tuset, Kevis-Kokitsi Maninis, and Pablo Arbelaez. In WACV 2025. [code]
-
SANPO: A Scene Understanding, Accessibility and Human Navigation Dataset - Sagar M. Waghmare, Kimberly Wilber, Dave Hawkey, Xuan Yang, Matthew Wilson, Stephanie Debats, et al. In WACV 2025. [project page] [code]
-
Trans4Map: Revisiting Holistic Bird's-Eye-View Mapping from Egocentric Images to Allocentric Semantics with Vision Transformers - Chang Chen, Jiaming Zhang, Kailun Yang, Kunyu Peng, and Rainer Stiefelhagen. In WACV 2023.
-
Intention-Conditioned Long-Term Human Egocentric Action Forecasting - Esteve Valls Mascaro, Hyemin Ahn, and Dongheui Lee. In WACV 2023.
-
Fine-grained Affordance Annotation for Egocentric Hand-Object Interaction Videos - Zecheng Yu, Yifei Huang, Ryosuke Furuta, Takuma Yagi, Yusuke Goutsu, and Yoichi Sato. In WACV 2023.
-
Domain Generalization through Audio-Visual Relative Norm Alignment in First Person Action Recognition - Mirco Planamente, Chiara Plizzari, Emanuele Alberti, and Barbara Caputo. In WACV 2022.
-
Integrating Human Gaze Into Attention for Egocentric Activity Recognition - Kyle Min, Jason J. Corso. In WACV 2021. [code]
-
Whose Hand Is This? Person Identification From Egocentric Hand Gestures - Satoshi Tsutsui, Yanwei Fu, and David J. Crandall. In WACV 2021.
-
The MECCANO Dataset: Understanding Human-Object Interactions from Egocentric Videos in an Industrial-like Domain - Francesco Ragusa, Antonino Furnari, Salvatore Livatino, and Giovanni Maria Farinella. In WACV 2021. [project page]
-
Automatic Calibration of the Fisheye Camera for Egocentric 3D Human Pose Estimation From a Single Image - Yahui Zhang, Shaodi You, and Theo Gevers. In WACV 2021.
-
Hand-Priming in Object Localization for Assistive Egocentric Vision - Kyungjun Lee, Abhinav Shrivastava, and Hernisa Kacorri. In WACV 2020.
-
Digging Deeper into Egocentric Gaze Prediction - Hamed R. Tavakoli, Esa Rahtu, Juho Kannala, and Ali Borji. In WACV 2019.
-
EGO-SLAM: A Robust Monocular SLAM for Egocentric Videos - Suvam Patra, Kartikeya Gupta, Faran Ahmad, Chetan Arora, and Subhashis Banerjee. In WACV 2019. [code]
-
Compact CNN for Indexing Egocentric Videos - Yair Poleg, Ariel Ephrat, Shmuel Peleg, and Chetan Arora. In WACV 2016.
Show papers (6)
-
Pandora: Articulated 3D Scene Graphs from Egocentric Vision - Alan Yu, Yun Chang, Christopher Xie, and Luca Carlone. In BMVC 2025.
-
Toward Robust Audio-Visual Synchronization Detection in Egocentric Video with Sparse Synchronization Events - Jordan Voas, Wei-Cheng Tseng, Benoit Vallade, Alex Mackin, David Higham, and David Harwath. In BMVC 2025.
-
With a Little Help from my Temporal Context: Multimodal Egocentric Action Recognition - Evangelos Kazakos, Jaesung Huh, Arsha Nagrani, Andrew Zisserman, and Dima Damen. In BMVC 2021. [project page] [code]
-
Stacked Temporal Attention: Improving First-person Action Recognition by Emphasizing Discriminative Clips - Lijin Yang, Yifei Huang, Yusuke Sugano, and Yoichi Sato. In BMVC 2021. [project page]
-
Hand-Object Contact Prediction via Motion-Based Pseudo-Labeling and Guided Progressive Label Correction - Takuma Yagi, Md Tasnimul Hasan, and Yoichi Sato. In BMVC 2021. [project page] [code]
-
You-Do, I-Learn: Discovering Task Relevant Objects and their Modes of Interaction from Multi-User Egocentric Video - Dima Damen, Tessid Leelasawassuk, Osian Haines, Andrew Calway,and Walterio Mayol-Cuevas. In BMVC 2014 [project page]
Show papers (38)
-
EgoDTM: Towards 3D-Aware Egocentric Video-Language Pretraining - Boshen Xu, Yuting Mei, Xinbi Liu, Sipeng Zheng, and Qin Jin. In NeurIPS 2025. [code]
-
EgoThinker: Unveiling Egocentric Reasoning with Spatio-Temporal CoT - Baoqi Pei, Yifei Huang, Jilan Xu, Yuping He, Guo Chen, Fei Wu, Yu Qiao, and Jiangmiao Pang. In NeurIPS 2025. [code]
-
EOC-Bench: Can MLLMs Identify, Recall, and Forecast Objects in an Egocentric World? - Yuqian Yuan, Ronghao Dang, Long Li, Wentong Li, Dian Jiao, Xin Li, Deli Zhao, Fan Wang, Wenqiao Zhang, Jun Xiao, and Yueting Zhuang. In NeurIPS 2025. [project page] [code]
-
WearVQA: A Visual Question Answering Benchmark for Wearables in Egocentric Authentic Real-world scenarios - Eun Chang, Zhuangqun Huang, Yiwei Liao, Sagar Ravi Bhavsar, Amogh Param, Tammy Stark, et al. In NeurIPS 2025.
-
Gaze-VLM: Bridging Gaze and VLMs through Attention Regularization for Egocentric Understanding - Anupam Pani and Yanchao Yang. In NeurIPS 2025. [code]
-
In the Eye of MLLM: Benchmarking Egocentric Video Intent Understanding with Gaze-Guided Prompting - Taiying Peng, Jiacheng Hua, Miao Liu, and Feng Lu. In NeurIPS 2025.
-
OpenMMEgo: Enhancing Egocentric Understanding for LMMs with Open Weights and Data - Hao Luo, Zihao Yue, Wanpeng Zhang, Yicheng Feng, Sipeng Zheng, Deheng Ye, and Zongqing Lu. In NeurIPS 2025. [code]
-
Eyes Wide Open: Ego Proactive Video-LLM for Streaming Video - Yulin Zhang, Cheng Shi, Yang Wang, and Sibei Yang. In NeurIPS 2025. [code]
-
PlayerOne: Egocentric World Simulator - Yuanpeng Tu, Hao Luo, Xi Chen, Xiang Bai, Fan Wang, and Hengshuang Zhao. In NeurIPS 2025. [project page] [code]
-
EgoVid-5M: A Large-Scale Video-Action Dataset for Egocentric Video Generation - Xiaofeng Wang, Kang Zhao, Feng Liu, Jiayu Wang, Guosheng Zhao, Xiaoyi Bao, Zheng Zhu, Yingya Zhang, and Xingang Wang. In NeurIPS 2025. [project page] [code]
-
EPFL-Smart-Kitchen: An Ego-Exo Multi-Modal Dataset for Challenging Action and Motion Understanding in Video-Language Models - Andy Bonnetto, Haozhe Qi, Franklin Leong, Matea Tashkovska, Mahdi Rad, Solaiman Shokur, Friedhelm Hummel, Silvestro Micera, Marc Pollefeys, and Alexander Mathis. In NeurIPS 2025. [project page] [code]
-
EgoExOR: An Ego-Exo-Centric Operating Room Dataset for Surgical Activity Understanding - Ege Özsoy, Arda Mamur, Felix Tristram, Chantal Pellegrini, Magdalena Wysocki, Benjamin Busam, and Nassir Navab. In NeurIPS 2025. [code]
-
Robust Ego-Exo Correspondence with Long-Term Memory - Yijun Hu, Bing Fan, Xin Gu, Haiqing Ren, Dongfang Liu, Heng Fan, and Libo Zhang. In NeurIPS 2025. [code]
-
IndEgo: A Dataset of Industrial Scenarios and Collaborative Work for Egocentric Assistants - Vivek Chavan, Yasmina Imgrund, Tung Dao, Sanwantri Bai, Bosong Wang, Ze Lu, Oliver Heimann, and Jörg Krüger. In NeurIPS 2025. [project page] [code] [dataset]
-
Seeing in the Dark: Benchmarking Egocentric 3D Vision with the Oxford Day-and-Night Dataset - Zirui Wang, Wenjing Bian, Xinghui Li, Yifu Tao, Jianeng Wang, Maurice Fallon, and Victor Adrian Prisacariu. In NeurIPS 2025. [project page]
-
egoEMOTION: Egocentric Vision and Physiological Signals for Emotion and Personality Recognition in Real-World Tasks - Matthias Jammot, Björn Braun, Paul Streli, Rafael Wampfler, and Christian Holz. In NeurIPS 2025. [project page] [code]
-
Gaze Beyond the Frame: Forecasting Egocentric 3D Visual Span - Heeseung Yun, Joonil Na, Jaeyeon Kim, Calvin Murdock, and Gunhee Kim. In NeurIPS 2025.
-
Robust Egocentric Referring Video Object Segmentation via Dual-Modal Causal Intervention - Haijing Liu, Zhiyuan Song, Hefeng Wu, Tao Pu, Keze Wang, and Liang Lin. In NeurIPS 2025.
-
Benchmarking Egocentric Multimodal Goal Inference for Assistive Wearable Agents - Vijay Veerabadran, Fanyi Xiao, Nitin Kamra, Pedro Matias, Joy Chen, Caley Drooff, et al. In NeurIPS 2025.
-
Whole-Body Conditioned Egocentric Video Prediction - Yutong Bai, Danny Tran, Amir Bar, Yann LeCun, Trevor Darrell, and Jitendra Malik. In NeurIPS 2025. [project page]
-
MEgoHand: Multimodal Egocentric Hand-Object Interaction Motion Generation - Bohan Zhou, Yi Zhan, Zhongbin Zhang, and Zongqing Lu. In NeurIPS 2025. [project page] [code]
-
EgoExoBench: A Benchmark for First- and Third-person View Video Understanding in MLLMs - Yuping He, Yifei Huang, Guo Chen, Baoqi Pei, Jilan Xu, Tong Lu, Jiangmiao Pang, et al. In NeurIPS 2025. [code]
-
EgoBlind: Towards Egocentric Visual Assistance for the Blind - Junbin Xiao, Nanxin Huang, Hao Qiu, Zhulin Tao, Xun Yang, Richang Hong, Meng Wang, and Angela Yao. In NeurIPS 2025.
-
EgoChoir: Capturing 3D Human-Object Interaction Regions from Egocentric Views - Yuhang Yang, Wei Zhai, Chengfeng Wang, Chengjun Yu, Yang Cao, and Zheng-Jun Zha. In NeurIPS 2024. [project page] [code]
-
HourVideo: 1-Hour Video-Language Understanding - Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu, and Li Fei-Fei. In NeurIPS 2024. [project page] [code]
-
CaptainCook4D: A Dataset for Understanding Errors in Procedural Activities - Rohith Peddi, Shivvrat Arya, Bharath Challa, Likhitha Pallapothula, Akshay Vyas, Bhavya Gouripeddi, et al. In NeurIPS 2024. [project page]
-
Streaming Detection of Queried Event Start - Cristóbal Eyzaguirre, Eric Tang, Shyamal Buch, Adrien Gaidon, Jiajun Wu, and Juan Carlos Niebles. In NeurIPS 2024. [project page]
-
Estimating Ego-Body Pose from Doubly Sparse Egocentric Video Data - Seunggeun Chi, Pin-Hao Huang, Enna Sachdeva, Hengbo Ma, Karthik Ramani, and Kwonjoon Lee. In NeurIPS 2024. [project page]
-
E³: Exploring Embodied Emotion Through A Large-Scale Egocentric Video Dataset - Wang Lin, Yueying Feng, Wenkang Han, Tao Jin, Zhou Zhao, Fei Wu, Chang Yao, and Jingyuan Chen. In NeurIPS 2024.
-
EgoSim: An Egocentric Multi-view Simulator and Real Dataset for Body-worn Cameras during Motion and Activity - Dominik Hollidt, Paul Streli, Jiaxi Jiang, Yasaman Haghighi, Changlin Qian, Xintong Liu, and Christian Holz. In NeurIPS 2024. [project page]
-
Exocentric-to-Egocentric Video Generation - Jia-Wei Liu, Weijia Mao, Zhongcong Xu, Jussi Keppo, and Mike Zheng Shou. In NeurIPS 2024. [code]
-
HENASY: Learning to Assemble Scene-Entities for Interpretable Egocentric Video-Language Model - Khoa Vo, Thinh Phan, Kashu Yamazaki, Minh Tran, and Ngan Le. In NeurIPS 2024. [project page] [code]
-
Differentiable Task Graph Learning: Procedural Activity Representation and Online Mistake Detection from Egocentric Videos - Luigi Seminara, Giovanni Maria Farinella, and Antonino Furnari. In NeurIPS 2024. [code]
-
EgoSchema: A Diagnostic Benchmark for Very Long-form Video Language Understanding - Karttikeya Mangalam, Raiymbek Akshulakov, and Jitendra Malik. In NeurIPS 2023. [project page] [code]
-
EgoTracks: A Long-term Egocentric Visual Object Tracking Dataset - Hao Tang, Kevin J Liang, Kristen Grauman, Matt Feiszli, and Weiyao Wang. In NeurIPS 2023. [dataset]
-
Egocentric Video-Language Pretraining - Kevin Qinghong Lin, Alex Jinpeng Wang, Mattia Soldan, Michael Wray, Rui Yan, Eric Zhongcong Xu, Difei Gao, Rongcheng Tu, Wenzhe Zhao, Weijie Kong, Chengfei Cai, Hongfa Wang, Dima Damen, Bernard Ghanem, Wei Liu and Mike Zheng Shou. In NeurIPS 2022. [project page] [code]
-
EgoTaskQA: Understanding Human Tasks in Egocentric Videos - Baoxiong Jia, Ting Lei, Song-Chun Zhu, Siyuan Huang. In NeurIPS 2022. [project page] [code]
-
Learning State-Aware Visual Representations from Audible Interactions - Himangi Mittal, Pedro Morgado, Unnat Jain, Abhinav Gupta. In NeurIPS 2022. [code] [Video]
A quick-reference table of some of the most prominent egocentric datasets/benchmarks. All entries below are also present in the full index in All Datasets; scale/modality/task figures are paraphrased from the descriptions already given there.
| Dataset | Scale | Modality/Sensors | Primary Task | Access |
|---|---|---|---|---|
| Ego4D | 3,025 hours, 855 wearers, 74 locations, 9 countries | Daily-life egocentric video | General-purpose egocentric benchmark suite | Requires Ego4D license |
| Ego-Exo4D | 1,286 hours, 740 participants, 13 cities | Multi-modal, multiview ego+exo video | Skilled human activity understanding (sports, music, dance, repair) | — |
| HOT3D | 833 minutes, 19 subjects, 33 objects | Multi-view egocentric image streams, 3D hand/object poses | Hand-object 3D tracking | — |
| Nymeria | 300h daily activity / 3,600h video, 1,200 sequences, 264 participants, 50 locations | Multiple egocentric multimodal devices | Human motion capture in the wild | — |
| Aria Digital Twin | 200 sequences, 398 object instances, 2 indoor scenes | Aria egocentric video + digital-twin ground truth | Egocentric 3D machine perception | — |
| HoloAssist | Large-scale (two-person sessions) | Egocentric human interaction video | Interactive AI assistants for physical manipulation tasks | — |
| HD-EPIC | 41 hours, 59.4K actions, 50.9K audio events, 26.6K VQA | Video, audio, 3D digital-twin grounding | Detailed kitchen action/audio understanding + VQA | — |
| EgoLife | ~300 hours, 6 participants, 1 week | Egocentric, interpersonal, multiview, multimodal (AI glasses) | Long-context daily-life assistance (EgoLifeQA) | — |
| EgoDex | 829 hours, 338K demonstrations, 194 tasks | Apple Vision Pro video + 3D head/hand pose + language | Tabletop manipulation demonstrations | — |
| EgoSchema | 5,000+ QA pairs, 250+ hours (from Ego4D) | Video question answering | Very long-form video-language understanding benchmark | — |
| Aria Everyday Activities | 143 sequences, 5 indoor locations | Project Aria (3D trajectories, point clouds, gaze, speech) | Daily-activity egocentric perception | — |
| EgoBody | Large-scale | Head-mounted device video | 3D human motion in social interactions | — |
| UnrealEgo | Large-scale | Egocentric stereo video | 3D human pose estimation | — |
| HOI4D | 2.4M RGB-D frames, 4,000 sequences, 9 participants, 800 objects, 16 categories | RGB-D egocentric video | Category-level human-object interaction | — |
| EgoExoLearn | 120 hours, 747 sequences | Ego+exo video with eye gaze | Procedural activity understanding across asynchronous views | — |
| CaptainCook4D | 384 recordings, 94.5 hours, 5.3K step / 10K action annotations | Egocentric 4D video | Procedure learning, error recognition | — |
| EgoProceL | 62 hours, 130 subjects, 16 tasks | Egocentric video | Procedure learning | — |
| MECCANO | 20 subjects | Egocentric video | Human-object interaction (industrial-like assembly) | — |
| EGTEA Gaze+ | 32 subjects, 86 sessions, 28 hours | Egocentric video + gaze | Cooking activity recognition and gaze | — |
| EPIC-Tent | 29 participants, dual head-mounted cameras | Egocentric video | Procedural activity (tent assembly) | — |
| IndEgo | 3,460 ego recordings (~197h) + 1,092 exo recordings (~97h) | Ego+exo video, gaze, narration, hand pose | Procedural task understanding, mistake detection, reasoning QA | — |
| EPFL-Smart-Kitchen-30 | 29.7 hours, 16 subjects, 4 recipes | Exo (9 RGB-D) + ego (HoloLens 2), depth, IMU, gaze, kinematics | Action and motion understanding benchmarks | — |
| HourVideo | 500 videos (20-120 min each), 12,976 QA pairs | Video question answering (from Ego4D) | Long-video language understanding | — |
- EgoHumans - 125K+ egocentric images from an in-the-wild multi-view multi-human capture setup (tennis, fencing, volleyball), with 3D pose, mesh, and tracking ground truth. [paper]
- EgoPet - About 84 hours of animal (dogs, cats, and others) egocentric video with interaction annotations, supporting benchmarks for visual interaction prediction, locomotion prediction, and vision-to-proprioception. [paper]
- Assembly101 - 4,321 videos of assembling/disassembling 101 take-apart toy vehicles, with 8 static and 4 egocentric views, 100K+ coarse and 1M fine-grained action segments, and 18M 3D hand poses, for procedural activity recognition, anticipation, segmentation, and mistake detection. [paper]
- EPIC-Contact - About 2,300 egocentric video clips (62,300 annotated frames) of bimanual hand-object interactions across nine everyday kitchen objects, for in-the-wild 3D hand-object pose estimation. [paper] [code]
- IndEgo - 3,460 egocentric recordings (~197 hours) plus 1,092 exocentric recordings (~97 hours) of industrial tasks including collaborative work, with eye gaze, narration, hand pose, mistake annotations, and benchmarks for procedural task understanding, mistake detection, and reasoning QA. [paper]
- Oxford Day-and-Night - Large-scale egocentric dataset captured with Meta Aria glasses spanning over 30 km of trajectories under day and night lighting, with multi-session SLAM ground-truth poses and 3D geometry for novel view synthesis and visual relocalisation benchmarks. [paper]
- egoEMOTION - Over 50 hours of recordings from 43 participants wearing Project Aria glasses, coupling eye-tracking video, head-mounted PPG, and inertial data with self-reported emotion and Big Five personality labels. [paper]
- LaMAria - City-scale egocentric visual-inertial SLAM benchmark captured with Aria glasses over hours and kilometers of trajectories, with survey-grade control points providing centimeter-accurate ground truth. [paper] [code]
- InterVLA - Large-scale egocentric human-object-human interaction dataset with 11.4 hours and 1.2M frames of multimodal data (2 egocentric and 5 exocentric views, human/object motion capture, and verbal commands). [paper]
- EPFL-Smart-Kitchen-30 - 29.7 hours of 16 subjects cooking four recipes with synchronized exocentric (9 RGB-D cameras) and egocentric (HoloLens 2) video, depth, IMUs, eye gaze, and body/hand kinematics, densely annotated for four action and motion understanding benchmarks. [paper]
- EgoExOR - 94 minutes (84,553 frames at 15 FPS) of two emulated spine procedures combining egocentric data (RGB, gaze, hand tracking, audio) from wearable glasses with exocentric RGB-D and ultrasound, annotated with 568,235 scene-graph triplets. [paper]
- EgoVid-5M - 5 million action-annotated egocentric video clips (1080p) with kinematic control signals and text descriptions, curated for egocentric video generation. [paper]
- EOC-Bench - A benchmark of 3,277 QA pairs over past/present/future temporal dimensions for evaluating object-centric embodied cognition of MLLMs in dynamic egocentric video. [paper]
- egoPPG-DB - 13+ hours of eye-tracking videos from Project Aria glasses of 25 participants performing everyday activities, with synchronized contact PPG and ECG-derived ground-truth heart rate. [paper]
- VISTA - A benchmark disentangling first-person viewpoint from the activity domain for visual object tracking and segmentation, with paired egocentric and third-person evaluation scenarios. [paper]
- EgoSchema - A very long-form video question-answering benchmark derived from Ego4D with over 5,000 human-curated multiple-choice QA pairs spanning over 250 hours of egocentric video, each question grounded in a three-minute clip. [paper] [code]
- EgoTracks - A long-term egocentric visual object tracking dataset built on Ego4D with 22.42k annotated tracks from 5.9k videos, released as part of the Ego4D benchmark; requires the Ego4D license for access. [paper]
- EgoBody3M - Large-scale real-image dataset for egocentric body tracking from VR-headset SLAM cameras, with more than 30 hours of recordings and about 3 million frames of diverse subjects and motions. [paper]
- Ego4D-Sounds - 1.2 million curated egocentric video clips from Ego4D with action-audio correspondence, built for action-to-sound generation. [paper]
- Ego4D-HCap - Hierarchical captioning dataset built on Ego4D with 8,267 manually annotated long-range video summaries for hour-long egocentric videos. [paper]
- EgoPER - 28 hours of egocentric procedural cooking videos across 5 tasks with normal and erroneous executions, multiple modalities (audio, depth, hand tracking), frame-wise step labels, and object bounding boxes for error detection. [paper] [code]
- HourVideo - A long-video benchmark of 500 manually curated egocentric videos from Ego4D (20 to 120 minutes each) with 12,976 five-way multiple-choice questions spanning summarization, perception, visual reasoning, and navigation tasks. [paper]
- CaptainCook4D - An egocentric 4D dataset of 384 recordings (94.5 hours) of people following or deviating from recipes in real kitchens, with 5.3K step and 10K fine-grained action annotations for error recognition, multistep localization, and procedure learning. [paper]
- Ego-EXTRA - 50 hours of unscripted egocentric videos of trainees performing procedural activities guided by real experts, with transcribed two-way dialogues, gaze, audio, and more than 15K visual question-answer sets for benchmarking egocentric video-language assistants. [paper]
- EgoXtreme - Egocentric 6D object pose estimation dataset of about 1.3M frames (775.5 minutes at 30 fps) captured with Project Aria glasses by 15 participants interacting with 13 objects under extreme lighting, motion blur, and smoke across industrial maintenance, sports, and emergency rescue scenarios. [paper]
- EgoAVU - Egocentric audio-visual understanding suite with a 3M-sample instruction-tuning set (EgoAVU-Instruct) and a manually verified evaluation benchmark (EgoAVU-Bench) covering grounding, temporal reasoning, scene understanding, and audio-visual hallucination. [paper]
- EgoEditData - Manually curated dataset of 100k egocentric video editing pairs featuring object substitution and removal under hand occlusions, interactions, and large egomotion. [paper]
- MyEgo - Egocentric VideoQA dataset with 541 long videos and 5K personalized questions about the camera wearer's things, activities, and past, designed to evaluate MLLM ego-grounding. [paper]
- EgoSound - Benchmark for egocentric sound understanding in MLLMs with 7,315 validated QA pairs across 900 videos sourced from Ego4D and EgoBlind, spanning a seven-task taxonomy from sound perception to cross-modal reasoning. [paper]
- Minerva-Ego - Egocentric video QA benchmark of 1,160 hand-crafted multiple-choice questions over 156 HD-EPIC videos, each paired with spatiotemporally grounded reasoning traces and object masks. [paper]
- Ego-1K - A large-scale multiview egocentric dataset of 956 short videos (~491K frames) captured by 12 synchronous cameras surrounding a VR headset at 1280×1280, for neural 3D video synthesis and dynamic scene understanding. [paper]
- Aria Gen 2 Pilot Dataset (A2PD) - Egocentric multimodal dataset captured with Aria Gen 2 glasses across everyday scenarios (cleaning, cooking, eating, playing, walking) between a primary user and three friends, with rich sensor streams (RGB, eye tracking, IMU, spatial audio, PPG, GPS) and machine-perception outputs.
- EgoDex - 829 hours of 30 fps 1080p egocentric video (338K demonstrations across 194 tabletop manipulation tasks) collected with Apple Vision Pro, paired with 3D head, upper-body, and hand pose plus natural-language annotations. [paper]
- EgoLife - A ~300-hour egocentric, interpersonal, multiview, multimodal dataset of six people living together for one week wearing AI glasses, accompanied by the EgoLifeQA long-context daily-assistance benchmark. [paper]
- HD-EPIC - 41 hours of highly-detailed unscripted kitchen recordings with 3D digital-twin grounding, 59.4K fine-grained actions, 50.9K audio events, 7.7M hand masks and 19.9K object tracks, plus a 26.6K-question VQA benchmark. [paper]
- EgoBlind - The first egocentric VideoQA dataset collected from blind/visually impaired users: 1,392 first-person videos and 5,311 questions posed or verified by blind individuals for assistive multimodal evaluation.
- EgoExoBench - A benchmark of 7,300+ QA pairs across 11 sub-tasks built from public ego-exo datasets to evaluate multimodal LLMs on cross-view (first- and third-person) video understanding.
- EgoPressure - An egocentric dataset of hand touch-contact and pressure interactions with accurate hand-pose meshes and per-contact pressure intensities.
- ParaHome - 486 minutes from 38 participants capturing 3D body and dexterous hand motion with multiple articulated household objects in a shared home environment, with text descriptions. [paper] [code]
- EgoMask - A pixel-level spatiotemporal grounding benchmark (with training set EgoMask-Train) built specifically for egocentric video. [paper]
- EgoExo-Fitness - Full-body action-understanding dataset of synchronized egocentric and exocentric fitness videos from 40 participants performing 86 types of fitness action sequences, with two-level temporal boundaries, technical-keypoint verification, language comments, and action-quality scores. [paper]
- EgoExoLearn - 120 hours across 747 sequences where camera wearers follow demonstrations to perform tasks in different environments, including eye-gaze signals, spanning daily food preparation to laboratory experiments. [paper]
- Aria Everyday Activities (AEA) - 143 daily-activity sequences recorded by multiple wearers in five indoor locations with Project Aria, with globally-aligned 3D trajectories, scene point clouds, per-frame eye gaze, and time-aligned speech. [paper]
- Ego4D Goal-Step - Hierarchical goal–step–substep annotations over Ego4D, with 2,807 hours carrying goal labels and 430 hours of fine-grained step labels (48K step segments). [paper]
- EventEgo3D - An egocentric event-stream dataset and benchmark for 3D human motion capture from a head-mounted event camera.
- ENIGMA-51 - An egocentric dataset of sequences recorded in an industrial laboratory, densely annotated for fine-grained human-object interaction understanding. [paper]
- EgoPoints - A benchmark for point tracking in egocentric video with challenging dynamic, occluded, and re-identification scenarios.
- SANPO - A human egocentric navigation dataset for scene understanding and accessibility, with real and synthetic first-person video and dense depth/segmentation.
- VIEW360 - An egocentric 360° video dataset for anomaly detection to assist people with visual impairments.
- Ego2HandsPose - An egocentric dataset for two-hand 3D global pose estimation.
- Aria Navigation Dataset (AND) - Roughly 4 hours of Project Aria egocentric recordings for humanoid/egocentric navigation, introduced with LookOut.
- MultiEgoView - 119 hours of synthetic (rendered from AMASS) plus 5 hours of real footage from 13 participants, captured from six body-worn cameras with full-body 3D pose ground truth; released with the EgoSim simulator. [paper]
- HOT3D - HOT3D is a dataset for benchmarking egocentric tracking of hands and objects in 3D. The dataset includes 833 minutes of multi-view image streams, which show 19 subjects interacting with 33 diverse rigid objects and are annotated with accurate 3D poses and shapes of hands and objects.
- Nymeria - Dataset of human motion in the wild, capturing diverse people engaging in diverse activities across diverse locations. Record body motion using multiple egocentric multimodal devices, all accurately synchronized and localized in one single metric 3D world. 300 hours of daily activity, 3600 hours of video data, 1200 sequences, 264 participants, 50 indoor and outdoor locations.
- Ego-Exo4D - 1,286 hours of multi-modal multiview videos recorded by 740 participants from 13 cities worldwide performing different skilled human activities (e.g., sports, music, dance, bike repair).
- Aria Digital Twin - A comprehensive egocentric dataset containing 200 sequences of real-world activities conducted by Aria wearers in two real indoor scenes with 398 object instances (324 stationary and 74 dynamic).
- HoloAssist - A large-scale egocentric human interaction dataset, where two people collaboratively complete physical manipulation tasks.
- EgoProceL - 62 hours of egocentric videos recorded by 130 subjects performing 16 tasks for procedure learning.
- EgoBody - Large-scale dataset capturing ground-truth 3D human motions during social interactions in 3D scenes.
- UnrealEgo - Large-scale naturalistic dataset for egocentric 3D human pose estimation.
- Hand-object Segments - Hand-object interactions in 11,235 frames from 1,000 videos covering daily activities in diverse scenarios.
- Ego4D - 3,025 hours of daily-life activity video spanning hundreds of scenarios (household, outdoor, workplace, leisure, etc.) captured by 855 unique camera wearers from 74 worldwide locations and 9 different countries.
- HOI4D - HOI4D consists of 2.4M RGB-D egocentric video frames over 4000 sequences collected by 9 participants interacting with 800 different object instances from 16 categories over 610 different indoor rooms.
- EgoCom - A natural conversations dataset containing multi-modal human communication data captured simultaneously from the participants' egocentric perspectives.
- TREK-100 - Object tracking in first person vision.
- MECCANO - 20 subject assembling a toy motorbike.
- EPIC-Kitchens 2020 - Subjects performing unscripted actions in their native environments.
- EPIC-Tent - 29 participants assembling a tent while wearing two head-mounted cameras. [paper]
- EGO-CH - 70 subjects visiting two cultural sites in Sicily, Italy.
- EPIC-Kitchens 2018 - 32 subjects performing unscripted actions in their native environments.
- Charade-Ego - Paired first-third person videos.
- EGTEA Gaze+ - 32 subjects, 86 cooking sessions, 28 hours.
- ADL - 20 subjects performing daily activities in their native environments.
- CMU kitchen - Multimodal, 18 subjects cooking 5 different recipes: brownies, eggs, pizza, salad, sandwich.
- EgoSeg - Long term actions (walking, running, driving, etc.)
- First-Person Social Interactions - 8 subjects at disneyworld.
- UEC Dataset - Two choreographed datasets with different egoactions (walk, jump, climb, etc.) + 6 YouTube sports videos.
- JPL - Interaction with a robot.
- FPPA - Five subjects performing 5 daily actions.
- UT Egocentric - 3-5 hours long videos capturing a person's day.
- VINST/ Visual Diaries - 31 videos capturing the visual experience of a subject walking from metro station to work.
- Bristol Egocentric Object Interaction (BEOID) - 8 subjects, six locations. Interaction with objects and environment.
- Object Search Dataset - 57 sequences of 55 subjects on search and retrieval tasks.
- UNICT-VEDI - Different subjects visiting a museum.
- UNICT-VEDI-POI - Different subjects visiting a museum.
- Simulated Egocentric Navigations - Simulated navigations of a virtual agent within a large building.
- EgoCart - Egocentric images collected by a shopping cart in a retail store.
- Unsupervised Segmentation of Daily Living Activities - Egocentric videos of daily activities.
- Visual Market Basket Analysis - Egocentric images collected by a shopping cart in a retail store.
- Location Based Segmentation of Egocentric Videos - Egocentric videos of daily activities.
- Recognition of Personal Locations from Egocentric Videos - Egocentric videos clips of daily.
- EgoGesture - 2k videos from 50 subjects performing 83 gestures.
- EgoHands - 48 videos of interactions between two people.
- DoMSEV - 80 hours/different activities.
- DR(eye)VE - 74 videos of people driving.
- THU-READ - 8 subjects performing 40 actions with a head-mounted RGBD camera.
- EgoDexter - 4 sequences with 4 actors (2 female), and varying interactions with various objects and and cluttered background. [paper]
- First-Person Hand Action (FPHA) - 3D hand-object interaction. Includes 1175 videos belonging to 45 different activity categories performed by 6 actors. [paper]
- UTokyo Paired Ego-Video (PEV) - 1,226 pairs of first-person clips extracted from the ones recorded synchronously during dyadic conversations.
- UTokyo Ego-Surf - Contains 8 diverse groups of first-person videos recorded synchronously during face-to-face conversations.
- TEgO: Teachable Egocentric Objects Dataset - Contains egocentric images of 19 distinct objects taken by two people for training a teachable object recognizer.
- Multimodal Focused Interaction Dataset - Contains 377 minutes of continuous multimodal recording captured during 19 sessions, with 17 conversational partners in 18 different indoor/outdoor locations.
Recurring workshops focused specifically on egocentric vision, useful for tracking the newest work between README updates:
- Joint Egocentric Vision (EgoVis) Workshop - The current main community venue for egocentric vision, merging the EPIC and Ego4D workshop series with Project Aria initiatives; held at CVPR 2024, 2025, and 2026, hosting challenges and a distinguished paper award.
- International Workshop on Egocentric Perception, Interaction and Computing (EPIC) - Long-running forum on egocentric perception across computer vision, machine learning, multimedia, AR/VR, and HCI, with editions at ECCV 2016, ICCV 2017, CVPR 2019, ICCV 2019, CVPR 2020, ECCV 2020, CVPR 2021, ICCV 2021, CVPR 2022, and CVPR 2023 (its final editions were joint with the Ego4D workshop; the series is now folded into EgoVis).
- International Ego4D Workshop - Challenge-centered workshop around the Ego4D benchmark suite, held at CVPR 2022 (joint with the 10th EPIC workshop), ECCV 2022, and CVPR 2023 (joint with the 11th EPIC workshop); now merged into EgoVis.
- EgoMotion: Egocentric Body Motion Tracking, Synthesis and Action Recognition - Workshop on human motion tracking, synthesis, and activity understanding from egocentric multimodal wearable-sensor data, held at CVPR 2024 and ICCV 2025.
This is a work in progress. Contributions welcome! Read the contribution guidelines first.