{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/visual-relationship-detection-with-internal","title":"Visual Relationship Detection with Internal and External Linguistic Knowledge Distillation","arxiv_id":"1707.09423","date":"2017-07-28","proceeding":"ICCV 2017 10","authors":["Ruichi Yu","Ang Li","Vlad I. Morariu","Larry S. Davis"],"abstract":"Understanding visual relationships involves identifying the subject, the\nobject, and a predicate relating them. We leverage the strong correlations\nbetween the predicate and the (subj,obj) pair (both semantically and spatially)\nto predict the predicates conditioned on the subjects and the objects. Modeling\nthe three entities jointly more accurately reflects their relationships, but\ncomplicates learning since the semantic space of visual relationships is huge\nand the training data is limited, especially for the long-tail relationships\nthat have few instances. To overcome this, we use knowledge of linguistic\nstatistics to regularize visual model learning. We obtain linguistic knowledge\nby mining from both training annotations (internal knowledge) and publicly\navailable text, e.g., Wikipedia (external knowledge), computing the conditional\nprobability distribution of a predicate given a (subj,obj) pair. Then, we\ndistill the knowledge into a deep model to achieve better generalization. Our\nexperimental results on the Visual Relationship Detection (VRD) and Visual\nGenome datasets suggest that with this linguistic knowledge distillation, our\nmodel outperforms the state-of-the-art methods significantly, especially when\npredicting unseen relationships (e.g., recall improved from 8.45% to 19.17% on\nVRD zero-shot testing set).","url_abs":"http://arxiv.org/abs/1707.09423v2","url_pdf":"http://arxiv.org/pdf/1707.09423v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"knowledge-distillation","task_name":"Knowledge Distillation"},{"task_slug":"relationship-detection","task_name":"Relationship Detection"},{"task_slug":"visual-relationship-detection","task_name":"Visual Relationship Detection"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/visual-relationship-detection-on-vrd-phrase","task":"Visual Relationship Detection","dataset":"VRD Phrase Detection","model":"Yu et. al [[Yu et al.2017a]]","rank_in_archive_order":1,"of":7,"metrics":{"R@100":"29.43","R@50":"26.32"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd","task":"Visual Relationship Detection","dataset":"VRD Predicate Detection","model":"Yu et. al [[Yu et al.2017a]]","rank_in_archive_order":1,"of":7,"metrics":{"R@100":"94.65","R@50":"85.64"},"uses_additional_data":false},{"leaderboard":"/sota/visual-relationship-detection-on-vrd-1","task":"Visual Relationship Detection","dataset":"VRD Relationship Detection","model":"Yu et. al [[Yu et al.2017a]]","rank_in_archive_order":1,"of":8,"metrics":{"R@100":"31.89","R@50":"22.68"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1707.09423","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}