{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-dynamic-graph-lstm-for-action-driven","title":"Temporal Dynamic Graph LSTM for Action-driven Video Object Detection","arxiv_id":"1708.00666","date":"2017-08-02","proceeding":"ICCV 2017 10","authors":["Yuan Yuan","Xiaodan Liang","Xiaolong Wang","Dit-yan Yeung","Abhinav Gupta"],"abstract":"In this paper, we investigate a weakly-supervised object detection framework.\nMost existing frameworks focus on using static images to learn object\ndetectors. However, these detectors often fail to generalize to videos because\nof the existing domain shift. Therefore, we investigate learning these\ndetectors directly from boring videos of daily activities. Instead of using\nbounding boxes, we explore the use of action descriptions as supervision since\nthey are relatively easy to gather. A common issue, however, is that objects of\ninterest that are not involved in human actions are often absent in global\naction descriptions known as \"missing label\". To tackle this problem, we\npropose a novel temporal dynamic graph Long Short-Term Memory network (TD-Graph\nLSTM). TD-Graph LSTM enables global temporal reasoning by constructing a\ndynamic graph that is based on temporal correlations of object proposals and\nspans the entire video. The missing label issue for each individual frame can\nthus be significantly alleviated by transferring knowledge across correlated\nobjects proposals in the whole video. Extensive evaluations on a large-scale\ndaily-life action dataset (i.e., Charades) demonstrates the superiority of our\nproposed method. We also release object bounding-box annotations for more than\n5,000 frames in Charades. We believe this annotated data can also benefit other\nresearch on video-based object recognition in the future.","url_abs":"http://arxiv.org/abs/1708.00666v1","url_pdf":"http://arxiv.org/pdf/1708.00666v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"object-recognition","task_name":"Object Recognition"},{"task_slug":"video-object-detection","task_name":"Video Object Detection"},{"task_slug":"weakly-supervised-object-detection","task_name":"Weakly Supervised Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[{"method_slug":"lstm","method_name":"LSTM"},{"method_slug":"memory-network","method_name":"Memory Network"},{"method_slug":"sigmoid-activation","method_name":"Sigmoid Activation"},{"method_slug":"tanh-activation","method_name":"Tanh Activation"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/weakly-supervised-object-detection-on-4","task":"Weakly Supervised Object Detection","dataset":"Charades","model":"TD-LSTM","rank_in_archive_order":3,"of":6,"metrics":{"MAP":"1.98"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.00666","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}