{"url":"/dataset/refcoco","name":"RefCOCO","full_name":null,"description_markdown":"The **RefCOCO dataset** is a **referring expression generation (REG)** dataset used for tasks related to understanding natural language expressions that refer to specific objects in images. Here are the key details about RefCOCO:\r\n\r\n1.  **Collection Method**:\r\nThe dataset was collected using the **ReferitGame**, a two-player game. In this game, the first player views an image with a segmented target object and writes a natural language expression referring to that object. The second player sees only the image and the referring expression and must click on the corresponding object. If both players perform correctly, they earn points and switch roles; otherwise, they receive a new object and image for description.\r\n\r\n2. **Dataset Variants**:\r\n**RefCOCO**: Contains 142,209 refer expressions for 50,000 objects across 19,994 images. **RefCOCO+**: Includes 141,564 expressions for 49,856 objects in 19,992 images. **RefCOCOg**: This variant has 25,799 images, 95,010 referring expressions, and 49,822 object instances.\r\n\r\n3. **Language and Restrictions**:\r\nRefCOCO allows **any type of language** in the referring expressions. RefCOCO+ disallows **location words** in expressions to focus purely on appearance-based descriptions (e.g., \"the man in the yellow polka-dotted shirt\") rather than viewer-dependent descriptions (e.g., \"the second man from the left\").\r\n\r\nThese datasets serve as valuable resources for tasks like referring expression segmentation, comprehension, and visual grounding in computer vision research.","description_withheld":null,"homepage":"https://github.com/lichengunc/refer","introduced_date":"2014-10-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/referitgame-referring-to-objects-in","title":"ReferItGame: Referring to Objects in Photographs of Natural Scenes","first_author":"Sahar Kazemzadeh","url":null},"license":{"name":"Unknown","url":null},"modalities":[{"name":"Images","url":"/datasets/modality/images"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Semantic Segmentation","url":"/task/semantic-segmentation","datasets_with_task":"/datasets/task/semantic-segmentation"},{"name":"Referring Expression Segmentation","url":"/task/referring-expression-segmentation","datasets_with_task":"/datasets/task/referring-expression-segmentation"},{"name":"Visual Reasoning","url":"/task/visual-reasoning","datasets_with_task":"/datasets/task/visual-reasoning"},{"name":"Referring Expression Comprehension","url":"/task/referring-expression-comprehension","datasets_with_task":"/datasets/task/referring-expression-comprehension"},{"name":"Visual Grounding","url":"/task/visual-grounding","datasets_with_task":"/datasets/task/visual-grounding"},{"name":"Zero-Shot Region Description","url":"/task/zero-shot-region-description","datasets_with_task":"/datasets/task/zero-shot-region-description"},{"name":"Region Proposal","url":"/task/region-proposal","datasets_with_task":"/datasets/task/region-proposal"}],"languages":[],"variants":["RefCOCOg-val","RefCOCOg-test","RefCoco+","RefCOCO+ test B","RefCOCO+ testA","RefCOCO+ val","RefCOCO testB","RefCOCO testA","RefCoCo val","RefCOCO"],"data_loaders":[{"repo":"https://github.com/tensorflow/datasets","url":"https://www.tensorflow.org/datasets/catalog/ref_coco","frameworks":["tf","jax"]},{"repo":"https://github.com/lichengunc/refer","url":"https://github.com/lichengunc/refer","frameworks":[]}],"num_papers_in_archive":439,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco","task":"Referring Expression Segmentation","dataset_variant":"RefCoCo val","rows":37,"metrics":["Overall IoU","mIoU","Precision@0.5","Precision@0.6","Precision@0.7","Precision@0.8","Precision@0.9","Mean IoU"],"first_row_in_archive_order":{"model":"DeRIS-L","paper":"/paper/deris-decoupling-perception-and-cognition-for","metrics":{"Mean IoU":"85.72","Overall IoU":"85.41"},"code_links":[{"title":"Dmmm1997/DeRIS","url":"https://github.com/Dmmm1997/DeRIS"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-3","task":"Referring Expression Segmentation","dataset_variant":"RefCOCO+ val","rows":33,"metrics":["Overall IoU","Mean IoU"],"first_row_in_archive_order":{"model":"MLCD-Seg-7B","paper":"/paper/multi-label-cluster-discrimination-for-visual","metrics":{"Overall IoU":"79.4"},"code_links":[{"title":"deepglint/unicom","url":"https://github.com/deepglint/unicom"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-4","task":"Referring Expression Segmentation","dataset_variant":"RefCOCO+ testA","rows":30,"metrics":["Overall IoU","Mean IoU","mIoU"],"first_row_in_archive_order":{"model":"HyperSeg","paper":"/paper/hyperseg-towards-universal-visual","metrics":{"Overall IoU":"83.5"},"code_links":[{"title":"congvvc/HyperSeg","url":"https://github.com/congvvc/HyperSeg"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-5","task":"Referring Expression Segmentation","dataset_variant":"RefCOCO+ test B","rows":30,"metrics":["Overall IoU","Mean IoU","mIoU"],"first_row_in_archive_order":{"model":"MLCD-Seg-7B","paper":"/paper/multi-label-cluster-discrimination-for-visual","metrics":{"Overall IoU":"75.6"},"code_links":[{"title":"deepglint/unicom","url":"https://github.com/deepglint/unicom"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-8","task":"Referring Expression Segmentation","dataset_variant":"RefCOCO testA","rows":13,"metrics":["Overall IoU","mIoU","Mean IoU"],"first_row_in_archive_order":{"model":"DeRIS-L","paper":"/paper/deris-decoupling-perception-and-cognition-for","metrics":{"Mean IoU":"86.64","Overall IoU":"86.49"},"code_links":[{"title":"Dmmm1997/DeRIS","url":"https://github.com/Dmmm1997/DeRIS"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-9","task":"Referring Expression Segmentation","dataset_variant":"RefCOCO testB","rows":13,"metrics":["Overall IoU","mIoU","Mean IoU"],"first_row_in_archive_order":{"model":"HyperSeg","paper":"/paper/hyperseg-towards-universal-visual","metrics":{"Overall IoU":"83.4"},"code_links":[{"title":"congvvc/HyperSeg","url":"https://github.com/congvvc/HyperSeg"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-grounding-on-refcoco-testa","task":"Visual Grounding","dataset_variant":"RefCOCO+ testA","rows":7,"metrics":["Accuracy (%)","IoU"],"first_row_in_archive_order":{"model":"Florence-2-large-ft","paper":"/paper/florence-2-advancing-a-unified-representation","metrics":{"Accuracy (%)":" 95.3"},"code_links":[{"title":"retkowsky/florence-2","url":"https://github.com/retkowsky/florence-2"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-grounding-on-refcoco-test-b","task":"Visual Grounding","dataset_variant":"RefCOCO+ test B","rows":6,"metrics":["Accuracy (%)"],"first_row_in_archive_order":{"model":"Florence-2-large-ft","paper":"/paper/florence-2-advancing-a-unified-representation","metrics":{"Accuracy (%)":"92.0"},"code_links":[{"title":"retkowsky/florence-2","url":"https://github.com/retkowsky/florence-2"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-grounding-on-refcoco-val","task":"Visual Grounding","dataset_variant":"RefCOCO+ val","rows":6,"metrics":["Accuracy (%)"],"first_row_in_archive_order":{"model":"Florence-2-large-ft","paper":"/paper/florence-2-advancing-a-unified-representation","metrics":{"Accuracy (%)":"93.4"},"code_links":[{"title":"retkowsky/florence-2","url":"https://github.com/retkowsky/florence-2"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-expression-segmentation-on-refcoco-6","task":"Referring Expression Segmentation","dataset_variant":"RefCOCO","rows":4,"metrics":["IoU","IoU (%)"],"first_row_in_archive_order":{"model":"DETRIS","paper":"/paper/densely-connected-parameter-efficient-tuning","metrics":{"IoU":"81.0"},"code_links":[{"title":"jiaqihuang01/detris","url":"https://github.com/jiaqihuang01/detris"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/visual-grounding-on-refcoco-testa-1","task":"Visual Grounding","dataset_variant":"RefCOCO testA","rows":1,"metrics":["IoU"],"first_row_in_archive_order":{"model":"HYDRA","paper":"/paper/hydra-a-hyper-agent-for-dynamic-compositional","metrics":{"IoU":"61.7"},"code_links":[{"title":"ControlNet/HYDRA","url":"https://github.com/ControlNet/HYDRA"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/deris-decoupling-perception-and-cognition-for","title":"DeRIS: Decoupling Perception and Cognition for Enhanced Referring Image Segmentation through Loopback Synergy","date":"2025-07-02","rows_on_this_dataset":6,"code_links":1,"syntology":null},{"paper":"/paper/segagent-exploring-pixel-understanding","title":"SegAgent: Exploring Pixel Understanding Capabilities in MLLMs by Imitating Human Annotator Trajectories","date":"2025-03-11","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/densely-connected-parameter-efficient-tuning","title":"Densely Connected Parameter-Efficient Tuning for Referring Image Segmentation","date":"2025-01-15","rows_on_this_dataset":7,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":7,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/multi-task-visual-grounding-with-coarse-to","title":"Multi-task Visual Grounding with Coarse-to-Fine Consistency Constraints","date":"2025-01-12","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":5,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/maskris-semantic-distortion-aware-data","title":"MaskRIS: Semantic Distortion-aware Data Augmentation for Referring Image Segmentation","date":"2024-11-28","rows_on_this_dataset":12,"code_links":1,"syntology":null},{"paper":"/paper/hyperseg-towards-universal-visual","title":"HyperSeg: Towards Universal Visual Segmentation with Large Language Model","date":"2024-11-26","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":7,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/multi-label-cluster-discrimination-for-visual","title":"Multi-label Cluster Discrimination for Visual Representation Learning","date":"2024-07-24","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":11,"samples_ran":7,"samples_unverified":4,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/safari-adaptive-sequence-transformer-for","title":"SafaRi:Adaptive Sequence Transformer for Weakly Supervised Referring Expression Segmentation","date":"2024-07-02","rows_on_this_dataset":6,"code_links":0,"syntology":null},{"paper":"/paper/evf-sam-early-vision-language-fusion-for-text","title":"EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model","date":"2024-06-28","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":3,"samples_ran":3,"samples_unverified":0,"pointer_only_for_licence":2,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/improving-referring-image-segmentation-using","title":"Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding","date":"2024-04-12","rows_on_this_dataset":6,"code_links":1,"syntology":null},{"paper":"/paper/psalm-pixelwise-segmentation-with-large-multi","title":"PSALM: Pixelwise SegmentAtion with Large Multi-Modal Model","date":"2024-03-21","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":2,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/hydra-a-hyper-agent-for-dynamic-compositional","title":"HYDRA: A Hyper Agent for Dynamic Compositional Visual Reasoning","date":"2024-03-19","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":5,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/groundhog-grounding-large-language-models-to","title":"GROUNDHOG: Grounding Large Language Models to Holistic Segmentation","date":"2024-02-26","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/mask-grounding-for-referring-image","title":"Mask Grounding for Referring Image Segmentation","date":"2023-12-19","rows_on_this_dataset":6,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":10,"samples_unverified":3,"pointer_only_for_licence":13,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/general-object-foundation-model-for-images","title":"General Object Foundation Model for Images and Videos at Scale","date":"2023-12-14","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":8,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/evp-enhanced-visual-perception-using-inverse","title":"EVP: Enhanced Visual Perception using Inverse Multi-Attentive Feature Refinement and Regularized Image-Text Alignment","date":"2023-12-13","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/universal-segmentation-at-arbitrary","title":"Universal Segmentation at Arbitrary Granularity with Language Instruction","date":"2023-12-04","rows_on_this_dataset":7,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":16,"samples_ran":13,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/florence-2-advancing-a-unified-representation","title":"Florence-2: Advancing a Unified Representation for a Variety of Vision Tasks","date":"2023-11-10","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/bridging-vision-and-language-encoders","title":"Bridging Vision and Language Encoders: Parameter-Efficient Tuning for Referring Image Segmentation","date":"2023-07-21","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":16,"samples_ran":8,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/hierarchical-open-vocabulary-universal-image-1","title":"Hierarchical Open-vocabulary Universal Image Segmentation","date":"2023-07-03","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/gres-generalized-referring-expression-1","title":"GRES: Generalized Referring Expression Segmentation","date":"2023-06-01","rows_on_this_dataset":4,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":2,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/universal-instance-perception-as-object","title":"Universal Instance Perception as Object Discovery and Retrieval","date":"2023-03-12","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/unleashing-text-to-image-diffusion-models-for-1","title":"Unleashing Text-to-Image Diffusion Models for Visual Perception","date":"2023-03-03","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":4,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/polyformer-referring-image-segmentation-as","title":"PolyFormer: Referring Image Segmentation as Sequential Polygon Generation","date":"2023-02-14","rows_on_this_dataset":8,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":5,"samples_unverified":1,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mplug-2-a-modularized-multi-modal-foundation","title":"mPLUG-2: A Modularized Multi-modal Foundation Model Across Text, Image and Video","date":"2023-02-01","rows_on_this_dataset":3,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":19,"samples_ran":9,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/toward-building-general-foundation-models-for","title":"Toward Building General Foundation Models for Language, Vision, and Vision-Language Understanding Tasks","date":"2023-01-12","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/x-2-vlm-all-in-one-pre-trained-model-for","title":"X$^2$-VLM: All-In-One Pre-trained Model For Vision-Language Tasks","date":"2022-11-22","rows_on_this_dataset":6,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":2,"samples_unverified":4,"pointer_only_for_licence":6,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/vlt-vision-language-transformer-and-query","title":"VLT: Vision-Language Transformer and Query Generation for Referring Segmentation","date":"2022-10-28","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/seqtr-a-simple-yet-universal-network-for","title":"SeqTR: A Simple yet Universal Network for Visual Grounding","date":"2022-03-30","rows_on_this_dataset":2,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":1,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/lavt-language-aware-vision-transformer-for","title":"LAVT: Language-Aware Vision Transformer for Referring Image Segmentation","date":"2021-12-04","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/cris-clip-driven-referring-image-segmentation","title":"CRIS: CLIP-Driven Referring Image Segmentation","date":"2021-11-30","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":12,"samples_ran":7,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/mail-a-unified-mask-image-language-trimodal","title":"MaIL: A Unified Mask-Image-Language Trimodal Network for Referring Image Segmentation","date":"2021-11-21","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/multi-grained-vision-language-pre-training","title":"Multi-Grained Vision Language Pre-Training: Aligning Texts with Visual Concepts","date":"2021-11-16","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/vision-language-transformer-and-query","title":"Vision-Language Transformer and Query Generation for Referring Segmentation","date":"2021-08-12","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/referring-transformer-a-one-step-approach-to","title":"Referring Transformer: A One-step Approach to Multi-task Visual Grounding","date":"2021-06-06","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":0,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/comprehensive-multi-modal-interactions-for","title":"Comprehensive Multi-Modal Interactions for Referring Image Segmentation","date":"2021-04-21","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/refvos-a-closer-look-at-referring-expressions","title":"RefVOS: A Closer Look at Referring Expressions for Video Object Segmentation","date":"2020-10-01","rows_on_this_dataset":5,"code_links":2,"syntology":null},{"paper":"/paper/referring-image-segmentation-via-cross-modal-1","title":"Referring Image Segmentation via Cross-Modal Progressive Comprehension","date":"2020-10-01","rows_on_this_dataset":4,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":0,"samples_unverified":7,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/bi-directional-relationship-inferring-network","title":"Bi-Directional Relationship Inferring Network for Referring Image Segmentation","date":"2020-06-01","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/referring-expression-object-segmentation-with","title":"Referring Expression Object Segmentation with Caption-Aware Consistency","date":"2019-10-10","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":2,"samples_ran":0,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/see-through-text-grouping-for-referring-image","title":"See-Through-Text Grouping for Referring Image Segmentation","date":"2019-10-01","rows_on_this_dataset":4,"code_links":0,"syntology":null},{"paper":"/paper/cross-modal-self-attention-network-for","title":"Cross-Modal Self-Attention Network for Referring Image Segmentation","date":"2019-04-09","rows_on_this_dataset":4,"code_links":1,"syntology":null},{"paper":"/paper/mattnet-modular-attention-network-for","title":"MAttNet: Modular Attention Network for Referring Expression Comprehension","date":"2018-01-24","rows_on_this_dataset":4,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":25,"samples_harvested":220,"samples_ran":111,"samples_unverified":109,"pointer_only_for_licence":27,"papers_with_no_sample_that_ran":4,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}