{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/moviegraphs-towards-understanding-human","title":"MovieGraphs: Towards Understanding Human-Centric Situations from Videos","arxiv_id":"1712.06761","date":"2017-12-19","proceeding":"CVPR 2018 6","authors":["Paul Vicol","Makarand Tapaswi","Lluis Castrejon","Sanja Fidler"],"abstract":"There is growing interest in artificial intelligence to build socially\nintelligent robots. This requires machines to have the ability to \"read\"\npeople's emotions, motivations, and other factors that affect behavior. Towards\nthis goal, we introduce a novel dataset called MovieGraphs which provides\ndetailed, graph-based annotations of social situations depicted in movie clips.\nEach graph consists of several types of nodes, to capture who is present in the\nclip, their emotional and physical attributes, their relationships (i.e.,\nparent/child), and the interactions between them. Most interactions are\nassociated with topics that provide additional details, and reasons that give\nmotivations for actions. In addition, most interactions and many attributes are\ngrounded in the video with time stamps. We provide a thorough analysis of our\ndataset, showing interesting common-sense correlations between different social\naspects of scenes, as well as across scenes over time. We propose a method for\nquerying videos and text with graphs, and show that: 1) our graphs contain rich\nand sufficient information to summarize and localize each scene; and 2)\nsubgraphs allow us to describe situations at an abstract level and retrieve\nmultiple semantically relevant situations. We also propose methods for\ninteraction understanding via ordering, and reason understanding. MovieGraphs\nis the first benchmark to focus on inferred properties of human-centric\nsituations, and opens up an exciting avenue towards socially-intelligent AI\nagents.","url_abs":"http://arxiv.org/abs/1712.06761v2","url_pdf":"http://arxiv.org/pdf/1712.06761v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"common-sense-reasoning","task_name":"Common Sense Reasoning"}],"methods":[],"datasets_introduced":[{"slug":"moviegraphs","name":"MovieGraphs","full_name":""}],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1712.06761","atlas_url":"https://app.syntology.ai/?focus=1712.06761","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}