{"url":"/dataset/segmentation-in-the-wild","name":"Segmentation in the Wild","full_name":"Segmentation in the Wild","description_markdown":"Recent advances in language-image pre-training has witnessed the emerging field of building transferable systems that can effortlessly adapt to a wide range of computer vision & multimodal tasks in the wild. This also poses a challenge to evaluate the transferability of these models due to the lack of easy-to-use evaluation toolkits and public benchmarks. \"Segmentation in the Wild (SegInW)\" Challenge is a part of X-Decoder, that proposed a new benchmark to evaluate the transfer ability of pre-trained vision models. This benchmark presents a diverse set of downstream segmentation datasets, measuring the ability of pre-training models on both the segmentation accuracy and their transfer efficiency in a new task, in terms of training examples and trainable parameters. This SegInW Challenge consists of 25 free public Segmentation datasets, crowd-sourced on roboflow.com. For more details about the challenge submission format, please refer to X-Decoder for SGinW.","description_withheld":null,"homepage":"https://eval.ai/web/challenges/challenge-page/1931/overview","introduced_date":"2022-12-21","introduced_date_note":null,"introduced_by":{"paper":"/paper/generalized-decoding-for-pixel-image-and","title":"Generalized Decoding for Pixel, Image, and Language","first_author":"Xueyan Zou","url":null},"license":{"name":"MIT license","url":"https://creativecommons.org/licenses/by-sa/4.0/"},"modalities":[],"tasks":[{"name":"Zero Shot Segmentation","url":"/task/zero-shot-segmentation","datasets_with_task":"/datasets/task/zero-shot-segmentation"}],"languages":[],"variants":["Segmentation in the Wild"],"data_loaders":[],"num_papers_in_archive":15,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/zero-shot-segmentation-on-segmentation-in-the","task":"Zero Shot Segmentation","dataset_variant":"Segmentation in the Wild","rows":12,"metrics":["Mean AP"],"first_row_in_archive_order":{"model":"Grounded HQ-SAM","paper":"/paper/segment-anything-in-high-quality","metrics":{"Mean AP":"49.6"},"code_links":[{"title":"huggingface/transformers","url":"https://github.com/huggingface/transformers"},{"title":"IDEA-Research/Grounded-Segment-Anything","url":"https://github.com/IDEA-Research/Grounded-Segment-Anything"},{"title":"syscv/sam-hq","url":"https://github.com/syscv/sam-hq"},{"title":"sqhuang0103/samreg","url":"https://github.com/sqhuang0103/samreg"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/opensd-unified-open-vocabulary-segmentation","title":"OpenSD: Unified Open-Vocabulary Segmentation and Detection","date":"2023-12-10","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/hierarchical-open-vocabulary-universal-image-1","title":"Hierarchical Open-vocabulary Universal Image Segmentation","date":"2023-07-03","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":2,"samples_unverified":2,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/segment-anything-in-high-quality","title":"Segment Anything in High Quality","date":"2023-06-02","rows_on_this_dataset":1,"code_links":4,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":3,"samples_unverified":14,"pointer_only_for_licence":3,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/a-simple-framework-for-open-vocabulary","title":"A Simple Framework for Open-Vocabulary Segmentation and Detection","date":"2023-03-14","rows_on_this_dataset":1,"code_links":2,"syntology":null},{"paper":"/paper/universal-instance-perception-as-object","title":"Universal Instance Perception as Object Discovery and Retrieval","date":"2023-03-12","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/grounding-dino-marrying-dino-with-grounded","title":"Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection","date":"2023-03-09","rows_on_this_dataset":1,"code_links":10,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":5,"samples_ran":2,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/open-vocabulary-panoptic-segmentation-with-1","title":"Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models","date":"2023-03-08","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/side-adapter-network-for-open-vocabulary","title":"Side Adapter Network for Open-Vocabulary Semantic Segmentation","date":"2023-02-23","rows_on_this_dataset":1,"code_links":3,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":1,"samples_unverified":8,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/generalized-decoding-for-pixel-image-and","title":"Generalized Decoding for Pixel, Image, and Language","date":"2022-12-21","rows_on_this_dataset":4,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":5,"samples_harvested":39,"samples_ran":11,"samples_unverified":28,"pointer_only_for_licence":3,"papers_with_no_sample_that_ran":0,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}