{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/actor-and-action-video-segmentation-from-a","title":"Actor and Action Video Segmentation from a Sentence","arxiv_id":"1803.07485","date":"2018-03-20","proceeding":"CVPR 2018 6","authors":["Kirill Gavrilyuk","Amir Ghodrati","Zhenyang Li","Cees G. M. Snoek"],"abstract":"This paper strives for pixel-level segmentation of actors and their actions\nin video content. Different from existing works, which all learn to segment\nfrom a fixed vocabulary of actor and action pairs, we infer the segmentation\nfrom a natural language input sentence. This allows to distinguish between\nfine-grained actors in the same super-category, identify actor and action\ninstances, and segment pairs that are outside of the actor and action\nvocabulary. We propose a fully-convolutional model for pixel-level actor and\naction segmentation using an encoder-decoder architecture optimized for video.\nTo show the potential of actor and action video segmentation from a sentence,\nwe extend two popular actor and action datasets with more than 7,500 natural\nlanguage descriptions. Experiments demonstrate the quality of the\nsentence-guided segmentations, the generalization ability of our model, and its\nadvantage for traditional actor and action segmentation compared to the\nstate-of-the-art.","url_abs":"http://arxiv.org/abs/1803.07485v1","url_pdf":"http://arxiv.org/pdf/1803.07485v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"actor-and-action-video-segmentation-from-a","repo_url":"https://github.com/JerryX1110/awesome-rvos","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"none","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"action-segmentation","task_name":"Action Segmentation"},{"task_slug":"decoder","task_name":"Decoder"},{"task_slug":"referring-expression-segmentation","task_name":"Referring Expression Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"sentence","task_name":"Sentence"},{"task_slug":"video-segmentation","task_name":"Video Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"a2d-sentences","name":"A2D Sentences","full_name":"Sentences for the Actor-Action Dataset (A2D)"}],"methods_introduced":[],"results":[{"leaderboard":"/sota/referring-expression-segmentation-on-a2d","task":"Referring Expression Segmentation","dataset":"A2D Sentences","model":"Gavriluyk el al. (Optical flow)","rank_in_archive_order":19,"of":27,"metrics":{"AP":"0.215","IoU mean":"0.426","IoU overall":"0.551","Precision@0.5":"0.5","Precision@0.6":"0.376","Precision@0.7":"0.231","Precision@0.8":"0.094","Precision@0.9":"0.004"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-a2d","task":"Referring Expression Segmentation","dataset":"A2D Sentences","model":"Gavriluyk el al.","rank_in_archive_order":20,"of":27,"metrics":{"AP":"0.198","IoU mean":"0.421","IoU overall":"0.536","Precision@0.5":"0.475","Precision@0.6":"0.347","Precision@0.7":"0.211","Precision@0.8":"0.08","Precision@0.9":"0.002"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-j-hmdb","task":"Referring Expression Segmentation","dataset":"J-HMDB","model":"Gavrilyuk et al. (Optical flow)","rank_in_archive_order":13,"of":21,"metrics":{"AP":"0.267","IoU mean":"0.570","IoU overall":"0.555","Precision@0.5":"0.712","Precision@0.6":"0.518","Precision@0.7":"0.264","Precision@0.8":"0.030","Precision@0.9":"0.000"},"uses_additional_data":false},{"leaderboard":"/sota/referring-expression-segmentation-on-j-hmdb","task":"Referring Expression Segmentation","dataset":"J-HMDB","model":"Gavrilyuk et al.","rank_in_archive_order":15,"of":21,"metrics":{"AP":"0.233","IoU mean":"0.542","IoU overall":"0.541","Precision@0.5":"0.699","Precision@0.6":"0.460","Precision@0.7":"0.173","Precision@0.8":"0.014","Precision@0.9":"0.000"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1803.07485","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}