{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/temporal-tessellation-a-unified-approach-for","title":"Temporal Tessellation: A Unified Approach for Video Analysis","arxiv_id":"1612.06950","date":"2016-12-21","proceeding":"ICCV 2017 10","authors":["Dotan Kaufman","Gil Levi","Tal Hassner","Lior Wolf"],"abstract":"We present a general approach to video understanding, inspired by semantic\ntransfer techniques that have been successfully used for 2D image analysis. Our\nmethod considers a video to be a 1D sequence of clips, each one associated with\nits own semantics. The nature of these semantics -- natural language captions\nor other labels -- depends on the task at hand. A test video is processed by\nforming correspondences between its clips and the clips of reference videos\nwith known semantics, following which, reference semantics can be transferred\nto the test video. We describe two matching methods, both designed to ensure\nthat (a) reference clips appear similar to test clips and (b), taken together,\nthe semantics of the selected reference clips is consistent and maintains\ntemporal coherence. We use our method for video captioning on the LSMDC'16\nbenchmark, video summarization on the SumMe and TVSum benchmarks, Temporal\nAction Detection on the Thumos2014 benchmark, and sound prediction on the\nGreatest Hits benchmark. Our method not only surpasses the state of the art, in\nfour out of five benchmarks, but importantly, it is the only single method we\nknow of that was successfully applied to such a diverse range of tasks.","url_abs":"http://arxiv.org/abs/1612.06950v2","url_pdf":"http://arxiv.org/pdf/1612.06950v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"temporal-tessellation-a-unified-approach-for","repo_url":"https://github.com/dot27/temporal-tessellation","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"tf","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"video-captioning","task_name":"Video Captioning"},{"task_slug":"video-summarization","task_name":"Video Summarization"},{"task_slug":"video-understanding","task_name":"Video Understanding"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-retrieval-on-msr-vtt","task":"Video Retrieval","dataset":"MSR-VTT","model":"Kaufman","rank_in_archive_order":39,"of":40,"metrics":{"text-to-video Median Rank":"41","text-to-video R@1":"4.7","text-to-video R@10":"24.1","video-to-text R@5":"16.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1612.06950","atlas_url":"https://app.syntology.ai/?focus=1612.06950","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}