{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semantic-embedding-space-for-zero-shot-action","title":"Semantic Embedding Space for Zero-Shot Action Recognition","arxiv_id":"1502.01540","date":"2015-02-05","proceeding":null,"authors":["Xun Xu","Timothy Hospedales","Shaogang Gong"],"abstract":"The number of categories for action recognition is growing rapidly. It is\nthus becoming increasingly hard to collect sufficient training data to learn\nconventional models for each category. This issue may be ameliorated by the\nincreasingly popular 'zero-shot learning' (ZSL) paradigm. In this framework a\nmapping is constructed between visual features and a human interpretable\nsemantic description of each category, allowing categories to be recognised in\nthe absence of any training data. Existing ZSL studies focus primarily on image\ndata, and attribute-based semantic representations. In this paper, we address\nzero-shot recognition in contemporary video action recognition tasks, using\nsemantic word vector space as the common space to embed videos and category\nlabels. This is more challenging because the mapping between the semantic space\nand space-time features of videos containing complex actions is more complex\nand harder to learn. We demonstrate that a simple self-training and data\naugmentation strategy can significantly improve the efficacy of this mapping.\nExperiments on human action datasets including HMDB51 and UCF101 demonstrate\nthat our approach achieves the state-of-the-art zero-shot action recognition\nperformance.","url_abs":"http://arxiv.org/abs/1502.01540v1","url_pdf":"http://arxiv.org/pdf/1502.01540v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"action-recognition-in-videos","task_name":"Action Recognition"},{"task_slug":"attribute","task_name":"Attribute"},{"task_slug":"data-augmentation","task_name":"Data Augmentation"},{"task_slug":"action-recognition","task_name":"Temporal Action Localization"},{"task_slug":"zero-shot-action-recognition","task_name":"Zero-Shot Action Recognition"},{"task_slug":"zero-shot-learning","task_name":"Zero-Shot Learning"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/zero-shot-action-recognition-on-ucf101","task":"Zero-Shot Action Recognition","dataset":"UCF101","model":"SVE","rank_in_archive_order":34,"of":35,"metrics":{"Top-1 Accuracy":"10.9"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1502.01540","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}