{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/commonly-uncommon-semantic-sparsity-in","title":"Commonly Uncommon: Semantic Sparsity in Situation Recognition","arxiv_id":"1612.00901","date":"2016-12-03","proceeding":"CVPR 2017 7","authors":["Mark Yatskar","Vicente Ordonez","Luke Zettlemoyer","Ali Farhadi"],"abstract":"Semantic sparsity is a common challenge in structured visual classification\nproblems; when the output space is complex, the vast majority of the possible\npredictions are rarely, if ever, seen in the training set. This paper studies\nsemantic sparsity in situation recognition, the task of producing structured\nsummaries of what is happening in images, including activities, objects and the\nroles objects play within the activity. For this problem, we find empirically\nthat most object-role combinations are rare, and current state-of-the-art\nmodels significantly underperform in this sparse data regime. We avoid many\nsuch errors by (1) introducing a novel tensor composition function that learns\nto share examples across role-noun combinations and (2) semantically augmenting\nour training data with automatically gathered examples of rarely observed\noutputs using web data. When integrated within a complete CRF-based structured\nprediction model, the tensor-based approach outperforms existing state of the\nart by a relative improvement of 2.11% and 4.40% on top-5 verb and noun-role\naccuracy, respectively. Adding 5 million images with our semantic augmentation\ntechniques gives further relative improvements of 6.23% and 9.57% on top-5 verb\nand noun-role accuracy.","url_abs":"http://arxiv.org/abs/1612.00901v1","url_pdf":"http://arxiv.org/pdf/1612.00901v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"commonly-uncommon-semantic-sparsity-in","repo_url":"https://github.com/my89/imSitu","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"commonly-uncommon-semantic-sparsity-in","repo_url":"https://github.com/thilinicooray/my_imsitu","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"grounded-situation-recognition","task_name":"Grounded Situation Recognition"},{"task_slug":"situation-recognition","task_name":"Situation Recognition"},{"task_slug":"structured-prediction","task_name":"Structured Prediction"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/grounded-situation-recognition-on-swig","task":"Grounded Situation Recognition","dataset":"SWiG","model":"CRF + Aug","rank_in_archive_order":12,"of":13,"metrics":{"Top-1 Verb":"34.12","Top-1 Verb & Value":"26.45","Top-5 Verbs":"62.59","Top-5 Verbs & Value":"46.88"},"uses_additional_data":false},{"leaderboard":"/sota/situation-recognition-on-imsitu","task":"Situation Recognition","dataset":"imSitu","model":"CRF + Aug","rank_in_archive_order":12,"of":13,"metrics":{"Top-1 Verb":"34.12","Top-1 Verb & Value":"26.45","Top-5 Verbs":"62.59","Top-5 Verbs & Value":"46.88"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/1612.00901","atlas_url":"https://app.syntology.ai/?focus=1612.00901","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}