{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/revisiting-the-effectiveness-of-off-the-shelf","title":"Revisiting the Effectiveness of Off-the-shelf Temporal Modeling Approaches for Large-scale Video Classification","arxiv_id":"1708.03805","date":"2017-08-12","proceeding":null,"authors":["Yunlong Bian","Chuang Gan","Xiao Liu","Fu Li","Xiang Long","Yandong Li","Heng Qi","Jie zhou","Shilei Wen","Yuanqing Lin"],"abstract":"This paper describes our solution for the video recognition task of\nActivityNet Kinetics challenge that ranked the 1st place. Most of existing\nstate-of-the-art video recognition approaches are in favor of an end-to-end\npipeline. One exception is the framework of DevNet. The merit of DevNet is that\nthey first use the video data to learn a network (i.e. fine-tuning or training\nfrom scratch). Instead of directly using the end-to-end classification scores\n(e.g. softmax scores), they extract the features from the learned network and\nthen fed them into the off-the-shelf machine learning models to conduct video\nclassification. However, the effectiveness of this line work has long-term been\nignored and underestimated. In this submission, we extensively use this\nstrategy. Particularly, we investigate four temporal modeling approaches using\nthe learned features: Multi-group Shifting Attention Network, Temporal Xception\nNetwork, Multi-stream sequence Model and Fast-Forward Sequence Model.\nExperiment results on the challenging Kinetics dataset demonstrate that our\nproposed temporal modeling approaches can significantly improve existing\napproaches in the large-scale video recognition tasks. Most remarkably, our\nbest single Multi-group Shifting Attention Network can achieve 77.7% in term of\ntop-1 accuracy and 93.2% in term of top-5 accuracy on the validation set.","url_abs":"http://arxiv.org/abs/1708.03805v1","url_pdf":"http://arxiv.org/pdf/1708.03805v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-classification","task_name":"Action Classification"},{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-classification","task_name":"Video Classification"},{"task_slug":"video-recognition","task_name":"Video Recognition"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-classification-on-kinetics-400","task":"Action Classification","dataset":"Kinetics-400","model":"Inception-ResNet","rank_in_archive_order":170,"of":207,"metrics":{"Acc@1":"73.0","Acc@5":"90.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1708.03805","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}