{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/egotaskqa-understanding-human-tasks-in","title":"EgoTaskQA: Understanding Human Tasks in Egocentric Videos","arxiv_id":"2210.03929","date":"2022-10-08","proceeding":null,"authors":["Baoxiong Jia","Ting Lei","Song-Chun Zhu","Siyuan Huang"],"abstract":"Understanding human tasks through video observations is an essential capability of intelligent agents. The challenges of such capability lie in the difficulty of generating a detailed understanding of situated actions, their effects on object states (i.e., state changes), and their causal dependencies. These challenges are further aggravated by the natural parallelism from multi-tasking and partial observations in multi-agent collaboration. Most prior works leverage action localization or future prediction as an indirect metric for evaluating such task understanding from videos. To make a direct evaluation, we introduce the EgoTaskQA benchmark that provides a single home for the crucial dimensions of task understanding through question-answering on real-world egocentric videos. We meticulously design questions that target the understanding of (1) action dependencies and effects, (2) intents and goals, and (3) agents' beliefs about others. These questions are divided into four types, including descriptive (what status?), predictive (what will?), explanatory (what caused?), and counterfactual (what if?) to provide diagnostic analyses on spatial, temporal, and causal understandings of goal-oriented tasks. We evaluate state-of-the-art video reasoning models on our benchmark and show their significant gaps between humans in understanding complex goal-oriented egocentric videos. We hope this effort will drive the vision community to move onward with goal-oriented video understanding and reasoning.","url_abs":"https://arxiv.org/abs/2210.03929v1","url_pdf":"https://arxiv.org/pdf/2210.03929v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"egotaskqa-understanding-human-tasks-in","repo_url":"https://github.com/Buzz-Beater/EgoTaskQA","is_official":1,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"action-localization","task_name":"Action Localization"},{"task_slug":"descriptive","task_name":"Descriptive"},{"task_slug":"diagnostic","task_name":"Diagnostic"},{"task_slug":"future-prediction","task_name":"Future prediction"},{"task_slug":"question-answering","task_name":"Question Answering"},{"task_slug":"video-understanding","task_name":"Video Understanding"},{"task_slug":null,"task_name":"counterfactual"}],"methods":[],"datasets_introduced":[{"slug":"egotaskqa","name":"EgoTaskQA","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2210.03929","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2210.03929"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Buzz-Beater/EgoTaskQA","reach":null}],"summary":{"ran_honours":1,"unverified":1},"by_repo_kind":{"official":{"samples":2,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"6ea690cb89faa443","entry":"reload","repo":"Buzz-Beater/EgoTaskQA","repo_kind":"official","path":"baselines/train_hcrn.py","file_url":"https://github.com/Buzz-Beater/EgoTaskQA/blob/HEAD/baselines/train_hcrn.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"well_formed","behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"6ea690cb89faa443"}},{"code_sha256_prefix":"7232d6166d6ad3d5","entry":"validate","repo":"Buzz-Beater/EgoTaskQA","repo_kind":"official","path":"baselines/train_hcrn.py","file_url":"https://github.com/Buzz-Beater/EgoTaskQA/blob/HEAD/baselines/train_hcrn.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"7232d6166d6ad3d5"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}