{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/discriminatively-learned-hierarchical-rank","title":"Discriminatively Learned Hierarchical Rank Pooling Networks","arxiv_id":"1705.10420","date":"2017-05-30","proceeding":null,"authors":["Basura Fernando","Stephen Gould"],"abstract":"In this work, we present novel temporal encoding methods for action and\nactivity classification by extending the unsupervised rank pooling temporal\nencoding method in two ways. First, we present \"discriminative rank pooling\" in\nwhich the shared weights of our video representation and the parameters of the\naction classifiers are estimated jointly for a given training dataset of\nlabelled vector sequences using a bilevel optimization formulation of the\nlearning problem. When the frame level features vectors are obtained from a\nconvolutional neural network (CNN), we rank pool the network activations and\njointly estimate all parameters of the model, including CNN filters and\nfully-connected weights, in an end-to-end manner which we coined as \"end-to-end\ntrainable rank pooled CNN\". Importantly, this model can make use of any\nexisting convolutional neural network architecture (e.g., AlexNet or VGG)\nwithout modification or introduction of additional parameters. Then, we extend\nrank pooling to a high capacity video representation, called \"hierarchical rank\npooling\". Hierarchical rank pooling consists of a network of rank pooling\nfunctions, which encode temporal semantics over arbitrary long video clips\nbased on rich frame level features. By stacking non-linear feature functions\nand temporal sub-sequence encoders one on top of the other, we build a high\ncapacity encoding network of the dynamic behaviour of the video. The resulting\nvideo representation is a fixed-length feature vector describing the entire\nvideo clip that can be used as input to standard machine learning classifiers.\nWe demonstrate our approach on the task of action and activity recognition.\nObtained results are comparable to state-of-the-art methods on three important\nactivity recognition benchmarks with classification performance of 76.7% mAP on\nHollywood2, 69.4% on HMDB51, and 93.6% on UCF101.","url_abs":"http://arxiv.org/abs/1705.10420v1","url_pdf":"http://arxiv.org/pdf/1705.10420v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"discriminatively-learned-hierarchical-rank","repo_url":"https://bitbucket.org/bfernando/hrp","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"activity-recognition","task_name":"Activity Recognition"},{"task_slug":"bilevel-optimization","task_name":"Bilevel Optimization"},{"task_slug":"classification","task_name":"General Classification"}],"methods":[{"method_slug":"1x1-convolution","method_name":"1x1 Convolution"},{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"dense-connections","method_name":"Dense Connections"},{"method_slug":"dropout","method_name":"Dropout"},{"method_slug":"grouped-convolution","method_name":"Grouped Convolution"},{"method_slug":"local-response-normalization","method_name":"Local Response Normalization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"relu","method_name":"ReLU"},{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":"https://syntology.ai/paper/1705.10420","atlas_url":"https://app.syntology.ai/?focus=1705.10420","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}