{"url":"/method/vc-r-cnn","slug":"vc-r-cnn","name":"VC R-CNN","full_name":"Visual Commonsense Region-based Convolutional Neural Network","full_name_withheld":false,"description_markdown":"**VC R-CNN** is an unsupervised feature representation learning method, which uses Region-based Convolutional Neural Network ([R-CNN](https://paperswithcode.com/method/r-cnn)) as the visual backbone, and the causal intervention as the training objective. Given a set of detected object regions in an image (e.g., using [Faster R-CNN](https://paperswithcode.com/method/faster-r-cnn)), like any other unsupervised feature learning methods (e.g., word2vec), the proxy training objective of VC R-CNN is to predict the contextual objects of a region. However, they are fundamentally different: the prediction of VC R-CNN is by using causal intervention: P(Y|do(X)), while others are by using the conventional likelihood: P(Y|X). This is also the core reason why VC R-CNN can learn \"sense-making\" knowledge like chair can be sat -- while not just \"common\" co-occurrences such as the chair is likely to exist if table is observed.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Visual Commonsense R-CNN","paper":"/paper/visual-commonsense-r-cnn","first_author":"Tan Wang","n_authors":4,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/visual-commonsense-r-cnn"},"source":{"url":"https://arxiv.org/abs/2002.12204v3","title":"Visual Commonsense R-CNN","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"General","area_id":"general","collection":"Self-Supervised Learning","url":"/methods/category/self-supervised-learning","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/visual-commonsense-r-cnn","title":"Visual Commonsense R-CNN","date":"2020-02-27","arxiv_id":"2002.12204","n_code_links":1,"syntology":{"ran":1,"of":4,"unverified":3,"pointer_only":0}}],"papers_shown":1,"tasks":[{"task":"/task/image-captioning","name":"Image Captioning","papers":1},{"task":"/task/representation-learning","name":"Representation Learning","papers":1},{"task":"/task/visual-question-answering","name":"Visual Question Answering (VQA)","papers":1}],"tasks_shown":3,"n_tasks":3,"usage_by_year":[{"year":"2020","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/vc-r-cnn"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}