{"url":"/dataset/refer-youtube-vos","name":"Refer-YouTube-VOS","full_name":null,"description_markdown":"There exist previous works [6, 10] that constructed referring segmentation datasets for videos. Gavrilyuk et al. [6] extended the A2D [33] and J-HMDB [9] datasets with natural sentences; the datasets focus on describing the ‘actors’ and ‘actions’ appearing in videos, therefore the instance annotations are limited to only a few object categories corresponding to the dominant ‘actors’ performing a salient ‘action’. Khoreva et al. [10] built a dataset based on DAVIS [25], but the scales are barely sufficient to learn an end-to-end model from scratch\r\n\r\nYoutube-VOS has 4,519 high-resolution videos with 94 common object categories. Each video has pixel-level instance segmentation annotation at every 5 frames in 30-fps videos, and their durations are around 3 to 6 seconds.\r\n\r\nWe employed Amazon Mechanical Turk to annotate referring expressions. To ensure the quality of the annotations, we selected around 50 turkers after a validation test. Each turker was given a pair of videos, the original video and the mask-overlaid one with the target object highlighted, and was asked to provide a discriminative sentence within 20 words that describes the target object accurately. We collected two kinds of annotations, which describe the highlighted object (1) based on a whole video (Full-video expression) and (2) using only the\r\nfirst frame of the video (First-frame expression). After the initial annotation, we conducted verification and cleaning jobs for all annotations, and dropped objects if an object cannot be localized using language expressions only. \r\n\r\nThe followings are the statistics and analysis of the two annotation types of the dataset after the verification.\r\n\r\n**Full-video expression:** Youtube-VOS has 6,459 and 1,063 unique objects in train and validation split, respectively. Among them, we cover 6,388 unique objects in 3,471 videos (6, 388/6, 459 = 98.9%) with 12,913 expressions in train split and 1,063 unique objects in 507 videos (1, 063/1, 063 = 100%) with 2,096 expressions in validation split. On average, each video has 3.8 language expressions and each expression has 10.0 words. \r\n\r\n**First-frame expression:** There are 6,006 unique objects in 3,412 videos (6, 006 /6, 459 = 93.0%) with 10,897 expressions in train split and 1,030 unique objects in 507 videos (1, 030/1, 063 = 96.9%) with 1,993 expressions in validation split. The number of annotated objects is lower than that of the full-video expressions because using only the first frame makes annotation more ambiguous and inconsistent and we dropped more annotations during the verification. On average,\r\neach video has 3.2 language expressions and each expression has 7.5 words.","description_withheld":null,"homepage":"https://youtube-vos.org/dataset/rvos/","introduced_date":"2020-08-01","introduced_date_note":null,"introduced_by":{"paper":"/paper/urvos-unified-referring-video-object","title":"URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark","first_author":"Seonguk Seo","url":null},"license":{"name":"Creative Commons Attribution 4.0 License","url":"https://youtube-vos.org/dataset/term/"},"modalities":[{"name":"Videos","url":"/datasets/modality/videos"},{"name":"Texts","url":"/datasets/modality/texts"}],"tasks":[{"name":"Referring Expression Segmentation","url":"/task/referring-expression-segmentation","datasets_with_task":"/datasets/task/referring-expression-segmentation"},{"name":"Referring Video Object Segmentation","url":"/task/referring-video-object-segmentation","datasets_with_task":"/datasets/task/referring-video-object-segmentation"}],"languages":[{"name":"English","url":"/datasets/language/english"}],"variants":["Refer-YouTube-VOS","Refer-YouTube-VOS (2021 public validation)"],"data_loaders":[],"num_papers_in_archive":52,"source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28"},"benchmarks":[{"leaderboard":"/sota/referring-expression-segmentation-on-refer-1","task":"Referring Expression Segmentation","dataset_variant":"Refer-YouTube-VOS (2021 public validation)","rows":33,"metrics":["J&F","J","F"],"first_row_in_archive_order":{"model":"MPG-SAM 2","paper":"/paper/mpg-sam-2-adapting-sam-2-with-mask-priors-and","metrics":{"F":"76.1","J":"71.7","J&F":"73.9"},"code_links":[{"title":"rongfu-dsb/MPG-SAM2","url":"https://github.com/rongfu-dsb/MPG-SAM2"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-video-object-segmentation-on-refer","task":"Referring Video Object Segmentation","dataset_variant":"Refer-YouTube-VOS","rows":18,"metrics":["J&F","J","F"],"first_row_in_archive_order":{"model":"FindTrack","paper":"/paper/find-first-track-next-decoupling","metrics":{"F":"75.7","J":"71.8","J&F":"73.7"},"code_links":[{"title":"suhwan-cho/FindTrack","url":"https://github.com/suhwan-cho/FindTrack"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"},{"leaderboard":"/sota/referring-expression-segmentation-on-refer","task":"Referring Expression Segmentation","dataset_variant":"Refer-YouTube-VOS","rows":2,"metrics":["Mean IoU","Precision@0.5","Precision@0.9"],"first_row_in_archive_order":{"model":"RefVOS-Human REs","paper":"/paper/synthref-generation-of-synthetic-referring","metrics":{"Mean IoU":"39.5","Precision@0.5":"38.6","Precision@0.9":"6.9"},"code_links":[{"title":"miriambellver/refvos","url":"https://github.com/miriambellver/refvos"},{"title":"imatge-upc/synthref","url":"https://github.com/imatge-upc/synthref"}]},"note":"rows are the archive's own order at snapshot; nothing here re-ranks them"}],"papers_with_a_benchmark_row":[{"paper":"/paper/find-first-track-next-decoupling","title":"Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation","date":"2025-03-05","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/referdino-referring-video-object-segmentation","title":"ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations","date":"2025-01-24","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/mpg-sam-2-adapting-sam-2-with-mask-priors-and","title":"MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation","date":"2025-01-23","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":16,"samples_ran":5,"samples_unverified":11,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/internvideo2-5-empowering-video-mllms-with","title":"InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling","date":"2025-01-21","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/the-devil-is-in-temporal-token-high-quality","title":"The Devil is in Temporal Token: High Quality Video Reasoning Segmentation","date":"2025-01-15","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/hyperseg-towards-universal-visual","title":"HyperSeg: Towards Universal Visual Segmentation with Large Language Model","date":"2024-11-26","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":17,"samples_ran":7,"samples_unverified":10,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/villa-video-reasoning-segmentation-with-large","title":"ViLLa: Video Reasoning Segmentation with Large Language Model","date":"2024-07-18","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/groprompt-efficient-grounded-prompting-and","title":"GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation","date":"2024-06-18","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/driving-referring-video-object-segmentation","title":"Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation","date":"2024-05-17","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/improving-referring-image-segmentation-using","title":"Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding","date":"2024-04-12","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/decoupling-static-and-hierarchical-motion","title":"Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation","date":"2024-04-04","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":2,"samples_unverified":5,"pointer_only_for_licence":7,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/towards-temporally-consistent-referring-video","title":"Temporally Consistent Referring Video Object Segmentation with Hybrid Memory","date":"2024-03-28","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":15,"samples_ran":14,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/univs-unified-and-universal-video","title":"UniVS: Unified and Universal Video Segmentation with Prompts as Queries","date":"2024-02-28","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":14,"samples_ran":12,"samples_unverified":2,"pointer_only_for_licence":14,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/uniref-segment-every-reference-object-in","title":"UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces","date":"2023-12-25","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":9,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/general-object-foundation-model-for-images","title":"General Object Foundation Model for Images and Videos at Scale","date":"2023-12-14","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":13,"samples_ran":8,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/universal-segmentation-at-arbitrary","title":"Universal Segmentation at Arbitrary Granularity with Language Instruction","date":"2023-12-04","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":16,"samples_ran":13,"samples_unverified":3,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/tracking-anything-with-decoupled-video","title":"Tracking Anything with Decoupled Video Segmentation","date":"2023-09-07","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":10,"samples_ran":7,"samples_unverified":3,"pointer_only_for_licence":10,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/epcformer-expression-prompt-collaboration","title":"Expression Prompt Collaboration Transformer for Universal Referring Video Object Segmentation","date":"2023-08-08","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/spectrum-guided-multi-granularity-referring","title":"Spectrum-guided Multi-granularity Referring Video Object Segmentation","date":"2023-07-25","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":9,"samples_ran":6,"samples_unverified":3,"pointer_only_for_licence":9,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/onlinerefer-a-simple-online-baseline-for","title":"OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation","date":"2023-07-18","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":7,"samples_ran":6,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/losh-long-short-text-joint-prediction-network","title":"LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation","date":"2023-06-14","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/soc-semantic-assisted-object-cluster-for","title":"SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation","date":"2023-05-26","rows_on_this_dataset":3,"code_links":1,"syntology":null},{"paper":"/paper/referred-by-multi-modality-a-unified-temporal","title":"Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation","date":"2023-05-25","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/universal-instance-perception-as-object","title":"Universal Instance Perception as Object Discovery and Retrieval","date":"2023-03-12","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":4,"samples_ran":3,"samples_unverified":1,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/segment-every-reference-object-in-spatial-and","title":"Segment Every Reference Object in Spatial and Temporal Spaces","date":"2023-01-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/html-hybrid-temporal-scale-multimodal","title":"HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object Segmentation","date":"2023-01-01","rows_on_this_dataset":6,"code_links":0,"syntology":null},{"paper":"/paper/vlt-vision-language-transformer-and-query","title":"VLT: Vision-Language Transformer and Query Generation for Referring Segmentation","date":"2022-10-28","rows_on_this_dataset":2,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":6,"samples_ran":0,"samples_unverified":6,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/multi-attention-network-for-compressed-video","title":"Multi-Attention Network for Compressed Video Referring Object Segmentation","date":"2022-07-26","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/r-2vos-robust-referring-video-object","title":"Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus","date":"2022-07-04","rows_on_this_dataset":2,"code_links":1,"syntology":null},{"paper":"/paper/deeply-interleaved-two-stream-encoder-for","title":"Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation","date":"2022-03-30","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/local-global-context-aware-transformer-for","title":"Local-Global Context Aware Transformer for Language-Guided Video Segmentation","date":"2022-03-18","rows_on_this_dataset":1,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":1,"samples_ran":1,"samples_unverified":0,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/language-as-queries-for-referring-video","title":"Language as Queries for Referring Video Object Segmentation","date":"2022-01-03","rows_on_this_dataset":3,"code_links":1,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":8,"samples_ran":7,"samples_unverified":1,"pointer_only_for_licence":8,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/multi-level-representation-learning-with","title":"Multi-Level Representation Learning With Semantic Alignment for Referring Video Object Segmentation","date":"2022-01-01","rows_on_this_dataset":1,"code_links":0,"syntology":null},{"paper":"/paper/end-to-end-referring-video-object","title":"End-to-End Referring Video Object Segmentation with Multimodal Transformers","date":"2021-11-29","rows_on_this_dataset":1,"code_links":2,"syntology":{"read_at":"2026-09-24T18:15:14+00:00","samples_harvested":11,"samples_ran":6,"samples_unverified":5,"pointer_only_for_licence":0,"claim":"Per-sample execution on synthesized fixtures; not a correctness claim."}},{"paper":"/paper/synthref-generation-of-synthetic-referring","title":"SynthRef: Generation of Synthetic Referring Expressions for Object Segmentation","date":"2021-06-08","rows_on_this_dataset":2,"code_links":2,"syntology":null},{"paper":"/paper/urvos-unified-referring-video-object","title":"URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark","date":"2020-08-01","rows_on_this_dataset":1,"code_links":1,"syntology":null},{"paper":"/paper/cross-modal-self-attention-network-for","title":"Cross-Modal Self-Attention Network for Referring Image Segmentation","date":"2019-04-09","rows_on_this_dataset":1,"code_links":1,"syntology":null}],"syntology_totals":{"read_at":"2026-09-24T18:15:14+00:00","papers_with_samples":16,"samples_harvested":163,"samples_ran":106,"samples_unverified":57,"pointer_only_for_licence":48,"papers_with_no_sample_that_ran":1,"note":"the per-paper counts above, summed; not a rate"},"papers_note":"The archive never published its papers-using-dataset list; these are papers with a leaderboard row on this dataset's benchmarks."}