{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/cuhk-ethz-siat-submission-to-activitynet","title":"CUHK & ETHZ & SIAT Submission to ActivityNet Challenge 2016","arxiv_id":"1608.00797","date":"2016-08-02","proceeding":null,"authors":["Yuanjun Xiong","Li-Min Wang","Zhe Wang","Bo-Wen Zhang","Hang Song","Wei Li","Dahua Lin","Yu Qiao","Luc van Gool","Xiaoou Tang"],"abstract":"This paper presents the method that underlies our submission to the untrimmed\nvideo classification task of ActivityNet Challenge 2016. We follow the basic\npipeline of temporal segment networks and further raise the performance via a\nnumber of other techniques. Specifically, we use the latest deep model\narchitecture, e.g., ResNet and Inception V3, and introduce new aggregation\nschemes (top-k and attention-weighted pooling). Additionally, we incorporate\nthe audio as a complementary channel, extracting relevant information via a CNN\napplied to the spectrograms. With these techniques, we derive an ensemble of\ndeep models, which, together, attains a high classification accuracy (mAP\n$93.23\\%$) on the testing set and secured the first place in the challenge.","url_abs":"http://arxiv.org/abs/1608.00797v1","url_pdf":"http://arxiv.org/pdf/1608.00797v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"cuhk-ethz-siat-submission-to-activitynet","repo_url":"https://github.com/yjxiong/anet2016-cuhk","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"classification","task_name":"General Classification"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"average-pooling","method_name":"Average Pooling"},{"method_slug":"batch-normalization","method_name":"Batch Normalization"},{"method_slug":"bottleneck-residual-block","method_name":"Bottleneck Residual Block"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"global-average-pooling","method_name":"Global Average Pooling"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"residual-block","method_name":"Residual Block"},{"method_slug":"residual-connection","method_name":"Residual Connection"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1608.00797","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}