{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/multi-granularity-generator-for-temporal","title":"Multi-granularity Generator for Temporal Action Proposal","arxiv_id":"1811.11524","date":"2018-11-28","proceeding":"CVPR 2019 6","authors":["Yuan Liu","Lin Ma","Yifeng Zhang","Wei Liu","Shih-Fu Chang"],"abstract":"Temporal action proposal generation is an important task, aiming to localize\nthe video segments containing human actions in an untrimmed video. In this\npaper, we propose a multi-granularity generator (MGG) to perform the temporal\naction proposal from different granularity perspectives, relying on the video\nvisual features equipped with the position embedding information. First, we\npropose to use a bilinear matching model to exploit the rich local information\nwithin the video sequence. Afterwards, two components, namely segment proposal\nproducer (SPP) and frame actionness producer (FAP), are combined to perform the\ntask of temporal action proposal at two distinct granularities. SPP considers\nthe whole video in the form of feature pyramid and generates segment proposals\nfrom one coarse perspective, while FAP carries out a finer actionness\nevaluation for each video frame. Our proposed MGG can be trained in an\nend-to-end fashion. By temporally adjusting the segment proposals with\nfine-grained frame actionness information, MGG achieves the superior\nperformance over state-of-the-art methods on the public THUMOS-14 and\nActivityNet-1.3 datasets. Moreover, we employ existing action classifiers to\nperform the classification of the proposals generated by MGG, leading to\nsignificant improvements compared against the competing methods for the video\ndetection task.","url_abs":"http://arxiv.org/abs/1811.11524v2","url_pdf":"http://arxiv.org/pdf/1811.11524v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"temporal-action-proposal-generation","task_name":"Temporal Action Proposal Generation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-thumos14","task":"Action Recognition","dataset":"THUMOS’14","model":"MGG UNet","rank_in_archive_order":2,"of":10,"metrics":{"mAP@0.3":"53.9","mAP@0.4":"46.8","mAP@0.5":"37.4"},"uses_additional_data":false},{"leaderboard":"/sota/temporal-action-proposal-generation-on","task":"Temporal Action Proposal Generation","dataset":"ActivityNet-1.3","model":"MGG","rank_in_archive_order":8,"of":11,"metrics":{"AR@100":"74.54","AUC (val)":"66.43"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1811.11524","atlas_url":"https://app.syntology.ai/?focus=1811.11524","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}