{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/rank-pooling-for-action-recognition","title":"Rank Pooling for Action Recognition","arxiv_id":"1512.01848","date":"2015-12-06","proceeding":null,"authors":["Basura Fernando","Efstratios Gavves","Jose Oramas","Amir Ghodrati","Tinne Tuytelaars"],"abstract":"We propose a function-based temporal pooling method that captures the latent\nstructure of the video sequence data - e.g. how frame-level features evolve\nover time in a video. We show how the parameters of a function that has been\nfit to the video data can serve as a robust new video representation. As a\nspecific example, we learn a pooling function via ranking machines. By learning\nto rank the frame-level features of a video in chronological order, we obtain a\nnew representation that captures the video-wide temporal dynamics of a video,\nsuitable for action recognition. Other than ranking functions, we explore\ndifferent parametric models that could also explain the temporal changes in\nvideos. The proposed functional pooling methods, and rank pooling in\nparticular, is easy to interpret and implement, fast to compute and effective\nin recognizing a wide variety of actions. We evaluate our method on various\nbenchmarks for generic action, fine-grained action and gesture recognition.\nResults show that rank pooling brings an absolute improvement of 7-10 average\npooling baseline. At the same time, rank pooling is compatible with and\ncomplementary to several appearance and local motion based methods and\nfeatures, such as improved trajectories and deep learning features.","url_abs":"http://arxiv.org/abs/1512.01848v2","url_pdf":"http://arxiv.org/pdf/1512.01848v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"rank-pooling-for-action-recognition","repo_url":"https://bitbucket.org/bfernando/videodarwin","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"none","reach":null}],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"gesture-recognition","task_name":"Gesture Recognition"},{"task_slug":"learning-to-rank","task_name":"Learning-To-Rank"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1512.01848","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}