{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/video-object-detection-with-an-aligned","title":"Video Object Detection with an Aligned Spatial-Temporal Memory","arxiv_id":"1712.06317","date":"2017-12-18","proceeding":"ECCV 2018 9","authors":["Fanyi Xiao","Yong Jae Lee"],"abstract":"We introduce Spatial-Temporal Memory Networks for video object detection. At\nits core, a novel Spatial-Temporal Memory module (STMM) serves as the recurrent\ncomputation unit to model long-term temporal appearance and motion dynamics.\nThe STMM's design enables full integration of pretrained backbone CNN weights,\nwhich we find to be critical for accurate detection. Furthermore, in order to\ntackle object motion in videos, we propose a novel MatchTrans module to align\nthe spatial-temporal memory from frame to frame. Our method produces\nstate-of-the-art results on the benchmark ImageNet VID dataset, and our\nablative studies clearly demonstrate the contribution of our different design\nchoices. We release our code and models at\nhttp://fanyix.cs.ucdavis.edu/project/stmn/project.html.","url_abs":"http://arxiv.org/abs/1712.06317v3","url_pdf":"http://arxiv.org/pdf/1712.06317v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"video-object-detection-with-an-aligned","repo_url":"https://github.com/fanyix/STMN","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":0,"framework":"torch","reach":{"status":"unanswered"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"object-detection","task_name":"Object Detection"},{"task_slug":"video-object-detection","task_name":"Video Object Detection"},{"task_slug":"object-detection-1","task_name":"object-detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1712.06317","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}