Datasets › Refer-YouTube-VOS

Refer-YouTube-VOS

Introduced by Seonguk Seo et al. in URVOS: Unified Referring Video Object Segmentation Network with a Large-Scale Benchmark1 Aug 2020 archive 2025-07-28

There exist previous works [6, 10] that constructed referring segmentation datasets for videos. Gavrilyuk et al. [6] extended the A2D [33] and J-HMDB [9] datasets with natural sentences; the datasets focus on describing the ‘actors’ and ‘actions’ appearing in videos, therefore the instance annotations are limited to only a few object categories corresponding to the dominant ‘actors’ performing a salient ‘action’. Khoreva et al. [10] built a dataset based on DAVIS [25], but the scales are barely sufficient to learn an end-to-end model from scratch

Youtube-VOS has 4,519 high-resolution videos with 94 common object categories. Each video has pixel-level instance segmentation annotation at every 5 frames in 30-fps videos, and their durations are around 3 to 6 seconds.

We employed Amazon Mechanical Turk to annotate referring expressions. To ensure the quality of the annotations, we selected around 50 turkers after a validation test. Each turker was given a pair of videos, the original video and the mask-overlaid one with the target object highlighted, and was asked to provide a discriminative sentence within 20 words that describes the target object accurately. We collected two kinds of annotations, which describe the highlighted object (1) based on a whole video (Full-video expression) and (2) using only the first frame of the video (First-frame expression). After the initial annotation, we conducted verification and cleaning jobs for all annotations, and dropped objects if an object cannot be localized using language expressions only.

The followings are the statistics and analysis of the two annotation types of the dataset after the verification.

Full-video expression: Youtube-VOS has 6,459 and 1,063 unique objects in train and validation split, respectively. Among them, we cover 6,388 unique objects in 3,471 videos (6, 388/6, 459 = 98.9%) with 12,913 expressions in train split and 1,063 unique objects in 507 videos (1, 063/1, 063 = 100%) with 2,096 expressions in validation split. On average, each video has 3.8 language expressions and each expression has 10.0 words.

First-frame expression: There are 6,006 unique objects in 3,412 videos (6, 006 /6, 459 = 93.0%) with 10,897 expressions in train split and 1,030 unique objects in 507 videos (1, 030/1, 063 = 96.9%) with 1,993 expressions in validation split. The number of annotated objects is lower than that of the full-video expressions because using only the first frame makes annotation more ambiguous and inconsistent and we dropped more annotations during the verification. On average, each video has 3.2 language expressions and each expression has 7.5 words.

Benchmarks archive 2025-07-28

All 3 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

Papers archive 2025-07-28

30 shown of 37 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 52. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Find First, Track Next: Decoupling Identification and Propagation in Referring Video Object Segmentation 1 1 5 Mar 2025 not harvested
ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations 0 1 24 Jan 2025 not harvested
MPG-SAM 2: Adapting SAM 2 with Mask Priors and Global Context for Referring Video Object Segmentation 1 1 23 Jan 2025 ran 5 of 16 samples (11 unverified)
InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling 1 1 21 Jan 2025 not harvested
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation 1 1 15 Jan 2025 not harvested
HyperSeg: Towards Universal Visual Segmentation with Large Language Model 1 1 26 Nov 2024 ran 7 of 17 samples (10 unverified)
ViLLa: Video Reasoning Segmentation with Large Language Model 1 1 18 Jul 2024 not harvested
GroPrompt: Efficient Grounded Prompting and Adaptation for Referring Video Object Segmentation 0 1 18 Jun 2024 not harvested
Harnessing Vision-Language Pretrained Models with Temporal-Aware Adaptation for Referring Video Object Segmentation 0 1 17 May 2024 not harvested
Vision-Aware Text Features in Referring Image Segmentation: From Object Understanding to Context Understanding 1 1 12 Apr 2024 not harvested
Decoupling Static and Hierarchical Motion Perception for Referring Video Segmentation 1 1 4 Apr 2024 ran 2 of 7 samples (5 unverified; 7 pointer-only for licence)
Temporally Consistent Referring Video Object Segmentation with Hybrid Memory 1 2 28 Mar 2024 ran 14 of 15 samples (1 unverified)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries 1 1 28 Feb 2024 ran 12 of 14 samples (2 unverified; 14 pointer-only for licence)
UniRef++: Segment Every Reference Object in Spatial and Temporal Spaces 2 1 25 Dec 2023 ran 9 of 9 samples (0 unverified)
General Object Foundation Model for Images and Videos at Scale 1 3 14 Dec 2023 ran 8 of 13 samples (5 unverified)
Universal Segmentation at Arbitrary Granularity with Language Instruction 2 1 4 Dec 2023 ran 13 of 16 samples (3 unverified)
Tracking Anything with Decoupled Video Segmentation 1 1 7 Sep 2023 ran 7 of 10 samples (3 unverified; 10 pointer-only for licence)
Expression Prompt Collaboration Transformer for Universal Referring Video Object Segmentation 0 1 8 Aug 2023 not harvested
Spectrum-guided Multi-granularity Referring Video Object Segmentation 1 2 25 Jul 2023 ran 6 of 9 samples (3 unverified; 9 pointer-only for licence)
OnlineRefer: A Simple Online Baseline for Referring Video Object Segmentation 1 1 18 Jul 2023 ran 6 of 7 samples (1 unverified)
LoSh: Long-Short Text Joint Prediction Network for Referring Video Object Segmentation 1 1 14 Jun 2023 not harvested
SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation 1 3 26 May 2023 not harvested
Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation 1 1 25 May 2023 not harvested
Universal Instance Perception as Object Discovery and Retrieval 1 1 12 Mar 2023 ran 3 of 4 samples (1 unverified)
Segment Every Reference Object in Spatial and Temporal Spaces 0 1 1 Jan 2023 not harvested
HTML: Hybrid Temporal-scale Multimodal Learning Framework for Referring Video Object Segmentation 0 6 1 Jan 2023 not harvested
VLT: Vision-Language Transformer and Query Generation for Referring Segmentation 1 2 28 Oct 2022 ran 0 of 6 samples (6 unverified)
Multi-Attention Network for Compressed Video Referring Object Segmentation 1 1 26 Jul 2022 not harvested
Towards Robust Referring Video Object Segmentation with Cyclic Relational Consensus 1 2 4 Jul 2022 not harvested
Deeply Interleaved Two-Stream Encoder for Referring Video Segmentation 0 1 30 Mar 2022 not harvested

The full list of 37 is in the JSON twin.

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons Attribution 4.0 License

Modalities archive 2025-07-28

Languages archive 2025-07-28

Variants archive 2025-07-28

  • Refer-YouTube-VOS
  • Refer-YouTube-VOS (2021 public validation)

2 variant names, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections