{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/untrimmed-video-classification-for-activity","title":"Untrimmed Video Classification for Activity Detection: submission to ActivityNet Challenge","arxiv_id":"1607.01979","date":"2016-07-07","proceeding":null,"authors":["Gurkirt Singh","Fabio Cuzzolin"],"abstract":"Current state-of-the-art human activity recognition is focused on the\nclassification of temporally trimmed videos in which only one action occurs per\nframe. We propose a simple, yet effective, method for the temporal detection of\nactivities in temporally untrimmed videos with the help of untrimmed\nclassification. Firstly, our model predicts the top k labels for each untrimmed\nvideo by analysing global video-level features. Secondly, frame-level binary\nclassification is combined with dynamic programming to generate the temporally\ntrimmed activity proposals. Finally, each proposal is assigned a label based on\nthe global label, and scored with the score of the temporal activity proposal\nand the global score. Ultimately, we show that untrimmed video classification\nmodels can be used as stepping stone for temporal detection.","url_abs":"http://arxiv.org/abs/1607.01979v2","url_pdf":"http://arxiv.org/pdf/1607.01979v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"untrimmed-video-classification-for-activity","repo_url":"https://github.com/gurkirt/actNet-inAct","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-detection","task_name":"Action Detection"},{"task_slug":"activity-detection","task_name":"Activity Detection"},{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"binary-classification","task_name":"Binary Classification"},{"task_slug":"classification-1","task_name":"Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"human-activity-recognition","task_name":"Human Activity Recognition"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1607.01979","atlas_url":"https://app.syntology.ai/?focus=1607.01979","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}