Papers › MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

MOMA-LRG: Language-Refined Graphs for Multi-Object Multi-Actor Activity Parsing

28 Nov 2022NeurIPS 2022 11archive 2025-07-28

Zelun Luo, Zane Durante, Linden Li, Wanze Xie, Ruochen Liu, Emily Jin, Zhuoyi Huang, Lun Yu Li, Jiajun Wu, Juan Carlos Niebles, Ehsan Adeli, Fei-Fei Li

Video-language models (VLMs), large models pre-trained on numerous but noisy video-text pairs from the internet, have revolutionized activity recognition through their remarkable generalization and open-vocabulary capabilities. While complex human activities are often hierarchical and compositional, most existing tasks for evaluating VLMs focus only on high-level video understanding, making it difficult to accurately assess and interpret the ability of VLMs to understand complex and fine-grained human activities. Inspired by the recently proposed MOMA framework, we define activity graphs as a single universal representation of human activities that encompasses video understanding at the activity, sub-activity, and atomic action level. We redefine activity parsing as the overarching task of activity graph generation, requiring understanding human activities across all three levels. To facilitate the evaluation of models on activity parsing, we introduce MOMA-LRG (Multi-Object Multi-Actor Language-Refined Graphs), a large dataset of complex human activities with activity graph annotations that can be readily transformed into natural language sentences. Lastly, we present a model-agnostic and lightweight approach to adapting and evaluating VLMs by incorporating structured knowledge from activity graphs into VLMs, addressing the individual limitations of language and graphical models. We demonstrate strong performance on few-shot activity parsing, and our framework is intended to foster future research in the joint modeling of videos, graphs, and language.

PaperPDFCode

Code

StanfordVL/moma mentioned in paper report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Activity RecognitionFew Shot Action RecognitionGraph GenerationVideo Understanding

Datasets

Introduced by this paper, per the archive.

MOMA-LRG

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Few Shot Action Recognition MOMA-LRG SG-VLM Activity Classification Accuracy (5-shot 5-way) 92.5 #2 of 4 Archive leaderboard report
Few Shot Action Recognition MOMA-LRG SG-VLM Subactivity Classification Accuracy (5-shot 5-way) 32.70 #2 of 4 Archive leaderboard report
Few Shot Action Recognition MOMA-LRG OTAM Activity Classification Accuracy (5-shot 5-way) 92.07 #3 of 4 Archive leaderboard report
Few Shot Action Recognition MOMA-LRG OTAM Subactivity Classification Accuracy (5-shot 5-way) 72.59 #3 of 4 Archive leaderboard report
Few Shot Action Recognition MOMA-LRG CMN Activity Classification Accuracy (5-shot 5-way) 86.3 #4 of 4 Archive leaderboard report
Few Shot Action Recognition MOMA-LRG CMN Subactivity Classification Accuracy (5-shot 5-way) 66.6 #4 of 4 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections