{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/long-tail-visual-relationship-recognition","title":"Exploring Long Tail Visual Relationship Recognition with Large Vocabulary","arxiv_id":"2004.00436","date":"2020-03-25","proceeding":"ICCV 2021 10","authors":["Sherif Abdelkarim","Aniket Agarwal","Panos Achlioptas","Jun Chen","Jiaji Huang","Boyang Li","Kenneth Church","Mohamed Elhoseiny"],"abstract":"Several approaches have been proposed in recent literature to alleviate the long-tail problem, mainly in object classification tasks. In this paper, we make the first large-scale study concerning the task of Long-Tail Visual Relationship Recognition (LTVRR). LTVRR aims at improving the learning of structured visual relationships that come from the long-tail (e.g., \"rabbit grazing on grass\"). In this setup, the subject, relation, and object classes each follow a long-tail distribution. To begin our study and make a future benchmark for the community, we introduce two LTVRR-related benchmarks, dubbed VG8K-LT and GQA-LT, built upon the widely used Visual Genome and GQA datasets. We use these benchmarks to study the performance of several state-of-the-art long-tail models on the LTVRR setup. Lastly, we propose a visiolinguistic hubless (VilHub) loss and a Mixup augmentation technique adapted to LTVRR setup, dubbed as RelMix. Both VilHub and RelMix can be easily integrated on top of existing models and despite being simple, our results show that they can remarkably improve the performance, especially on tail classes. Benchmarks, code, and models have been made available at: https://github.com/Vision-CAIR/LTVRR.","url_abs":"https://arxiv.org/abs/2004.00436v7","url_pdf":"https://arxiv.org/pdf/2004.00436v7.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"long-tail-visual-relationship-recognition","repo_url":"https://github.com/Vision-CAIR/LTVRR","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":null},{"paper_slug":"long-tail-visual-relationship-recognition","repo_url":"https://github.com/Elhoseiny-VisionCAIR-Lab/LTVRR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}},{"paper_slug":"long-tail-visual-relationship-recognition","repo_url":"https://github.com/sherif-abdelkarim/LTVRR","is_official":0,"mentioned_in_paper":0,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok","spdx":"MIT"}}],"tasks":[{"task_slug":"visual-relationship-detection","task_name":"Visual Relationship Detection"}],"methods":[{"method_slug":"mixup","method_name":"Mixup"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=2004.00436","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"2004.00436"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Elhoseiny-VisionCAIR-Lab/LTVRR","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/sherif-abdelkarim/LTVRR","reach":{"status":"ok","spdx":"MIT"}},{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/Vision-CAIR/LTVRR","reach":null}],"summary":{"ran_draft_wrong":1,"ran_honours":1,"unverified":3},"by_repo_kind":{"official":{"samples":2,"ran":2,"repositories":1},"listed":{"samples":3,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":0,"samples":[{"code_sha256_prefix":"83ac5f312d7bb313","entry":"create_one_hot","repo":"Vision-CAIR/LTVRR","repo_kind":"official","path":"lib/modeling/reldn_heads.py","file_url":"https://github.com/Vision-CAIR/LTVRR/blob/HEAD/lib/modeling/reldn_heads.py","link_basis":"first_harvest_node","language":"python","status":"ran_draft_wrong","verification_level":1,"contract_check":"MISDECLARED","metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"83ac5f312d7bb313"}},{"code_sha256_prefix":"23c77da2f0afb733","entry":"manual_CE","repo":"Vision-CAIR/LTVRR","repo_kind":"official","path":"lib/modeling/reldn_heads.py","file_url":"https://github.com/Vision-CAIR/LTVRR/blob/HEAD/lib/modeling/reldn_heads.py","link_basis":"first_harvest_node","language":"python","status":"ran_honours","verification_level":1,"contract_check":"HONOURS","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"23c77da2f0afb733"}},{"code_sha256_prefix":"7d4002289969a288","entry":"argsort_desc","repo":"Elhoseiny-VisionCAIR-Lab/LTVRR","repo_kind":"listed","path":"lib/datasets/task_evaluation_rel.py","file_url":"https://github.com/Elhoseiny-VisionCAIR-Lab/LTVRR/blob/HEAD/lib/datasets/task_evaluation_rel.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"7d4002289969a288"}},{"code_sha256_prefix":"6727d2d8f1d79998","entry":"intersect_2d","repo":"Elhoseiny-VisionCAIR-Lab/LTVRR","repo_kind":"listed","path":"lib/datasets/task_evaluation_rel.py","file_url":"https://github.com/Elhoseiny-VisionCAIR-Lab/LTVRR/blob/HEAD/lib/datasets/task_evaluation_rel.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"6727d2d8f1d79998"}},{"code_sha256_prefix":"f7051976b9be472b","entry":"voc_ap","repo":"Elhoseiny-VisionCAIR-Lab/LTVRR","repo_kind":"listed","path":"lib/datasets/voc_eval_rel.py","file_url":"https://github.com/Elhoseiny-VisionCAIR-Lab/LTVRR/blob/HEAD/lib/datasets/voc_eval_rel.py","link_basis":"harvester_set","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"MIT","inline_ok":true,"mcp_get_code":{"code_sha256":"f7051976b9be472b"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}