{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/fine-grained-video-categorization-with","title":"Fine-grained Video Categorization with Redundancy Reduction Attention","arxiv_id":"1810.11189","date":"2018-10-26","proceeding":"ECCV 2018 9","authors":["Chen Zhu","Xiao Tan","Feng Zhou","Xiao Liu","Kaiyu Yue","Errui Ding","Yi Ma"],"abstract":"For fine-grained categorization tasks, videos could serve as a better source\nthan static images as videos have a higher chance of containing discriminative\npatterns. Nevertheless, a video sequence could also contain a lot of redundant\nand irrelevant frames. How to locate critical information of interest is a\nchallenging task. In this paper, we propose a new network structure, known as\nRedundancy Reduction Attention (RRA), which learns to focus on multiple\ndiscriminative patterns by sup- pressing redundant feature channels.\nSpecifically, it firstly summarizes the video by weight-summing all feature\nvectors in the feature maps of selected frames with a spatio-temporal soft\nattention, and then predicts which channels to suppress or to enhance according\nto this summary with a learned non-linear transform. Suppression is achieved by\nmodulating the feature maps and threshing out weak activations. The updated\nfeature maps are then used in the next iteration. Finally, the video is\nclassified based on multiple summaries. The proposed method achieves out-\nstanding performances in multiple video classification datasets. Further- more,\nwe have collected two large-scale video datasets, YouTube-Birds and\nYouTube-Cars, for future researches on fine-grained video categorization. The\ndatasets are available at http://www.cs.umd.edu/~chenzhu/fgvc.","url_abs":"http://arxiv.org/abs/1810.11189v1","url_pdf":"http://arxiv.org/pdf/1810.11189v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"video-classification","task_name":"Video Classification"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/action-recognition-in-videos-on-activitynet","task":"Action Recognition","dataset":"ActivityNet","model":"RRA","rank_in_archive_order":12,"of":16,"metrics":{"mAP":"83.4"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1810.11189","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}