{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/the-pros-and-cons-rank-aware-temporal","title":"The Pros and Cons: Rank-aware Temporal Attention for Skill Determination in Long Videos","arxiv_id":"1812.05538","date":"2018-12-13","proceeding":"CVPR 2019 6","authors":["Hazel Doughty","Walterio Mayol-Cuevas","Dima Damen"],"abstract":"We present a new model to determine relative skill from long videos, through\nlearnable temporal attention modules. Skill determination is formulated as a\nranking problem, making it suitable for common and generic tasks. However, for\nlong videos, parts of the video are irrelevant for assessing skill, and there\nmay be variability in the skill exhibited throughout a video. We therefore\npropose a method which assesses the relative overall level of skill in a long\nvideo by attending to its skill-relevant parts. Our approach trains temporal\nattention modules, learned with only video-level supervision, using a novel\nrank-aware loss function. In addition to attending to task relevant video\nparts, our proposed loss jointly trains two attention modules to separately\nattend to video parts which are indicative of higher (pros) and lower (cons)\nskill. We evaluate our approach on the EPIC-Skills dataset and additionally\nannotate a larger dataset from YouTube videos for skill determination with five\npreviously unexplored tasks. Our method outperforms previous approaches and\nclassic softmax attention on both datasets by over 4% pairwise accuracy, and as\nmuch as 12% on individual tasks. We also demonstrate our model's ability to\nattend to rank-aware parts of the video.","url_abs":"http://arxiv.org/abs/1812.05538v2","url_pdf":"http://arxiv.org/pdf/1812.05538v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"the-pros-and-cons-rank-aware-temporal","repo_url":"https://github.com/hazeld/rank-aware-attention-network","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[],"methods":[{"method_slug":"softmax","method_name":"Softmax"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=1812.05538","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}