{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/shuffle-then-assemble-learning-object","title":"Shuffle-Then-Assemble: Learning Object-Agnostic Visual Relationship Features","arxiv_id":"1808.00171","date":"2018-08-01","proceeding":"ECCV 2018 9","authors":["Xu Yang","Hanwang Zhang","Jianfei Cai"],"abstract":"Due to the fact that it is prohibitively expensive to completely annotate\nvisual relationships, i.e., the (obj1, rel, obj2) triplets, relationship models\nare inevitably biased to object classes of limited pairwise patterns, leading\nto poor generalization to rare or unseen object combinations. Therefore, we are\ninterested in learning object-agnostic visual features for more generalizable\nrelationship models. By \"agnostic\", we mean that the feature is less likely\nbiased to the classes of paired objects. To alleviate the bias, we propose a\nnovel \\texttt{Shuffle-Then-Assemble} pre-training strategy. First, we discard\nall the triplet relationship annotations in an image, leaving two unpaired\nobject domains without obj1-obj2 alignment. Then, our feature learning is to\nrecover possible obj1-obj2 pairs. In particular, we design a cycle of residual\ntransformations between the two domains, to capture shared but not\nobject-specific visual patterns. Extensive experiments on two visual\nrelationship benchmarks show that by using our pre-trained features, naive\nrelationship models can be consistently improved and even outperform other\nstate-of-the-art relationship models. Code has been made available at:\n\\url{https://github.com/yangxuntu/vrd}.","url_abs":"http://arxiv.org/abs/1808.00171v1","url_pdf":"http://arxiv.org/pdf/1808.00171v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"shuffle-then-assemble-learning-object","repo_url":"https://github.com/yangxuntu/vrd","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":{"status":"ok"}}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":null,"task_name":"Triplet"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.00171","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}