Papers › Jointly Learning Energy Expenditures and Activities Using Egocentric Multimodal Signals

Jointly Learning Energy Expenditures and Activities Using Egocentric Multimodal Signals

1 Jul 2017CVPR 2017 7archive 2025-07-28

Katsuyuki Nakamura, Serena Yeung, Alexandre Alahi, Li Fei-Fei

Physiological signals such as heart rate can provide valuable information about an individual's state and activity. However, existing work on computer vision has not yet explored leveraging these signals to enhance egocentric video understanding. In this work, we propose a model for reasoning on multimodal data to jointly predict activities and energy expenditures. We use heart rate signals as privileged self-supervision to derive energy expenditure in a training stage. A multitask objective is used to jointly optimize the two tasks. Additionally, we introduce a dataset that contains 31 hours of egocentric video augmented with heart rate and acceleration signals. This study can lead to new applications such as a visual calorie counter.

PaperPDF

Code

No code repository is listed for this paper in the archive or in Syntology's graph.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Video Understanding

Datasets

Introduced by this paper, per the archive.

Stanford-ECM

Results from the paper archive 2025-07-28

No leaderboard rows for this paper in the archive.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections