Datasets › ReVOS

ReVOS

Introduced by Cilin Yan et al. in VISA: Reasoning Video Object Segmentation via Large Language Models16 Jul 2024 archive 2025-07-28

We create a benchmark dataset named ReVOS. This dataset comprises 35,074 pairs of instruction-mask sequences derived from 1,042 diverse videos. In contrast to traditional referring video segmentation datasets, such as Ref-YouTube-VOS and MeViS, which primarily contain explicit short phrases, ReVOS includes text instructions that necessitates a sophisticated understanding of both video content and general world knowledge

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Referring Video Object Segmentation ReVOS VRS-HQ (Chat-UniVi-13B) J&F 60 The Devil is in Temporal Token: High Quality Video... sitonggong/vrs-hq 9 Compare

Papers archive 2025-07-28

7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 18. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
The Devil is in Temporal Token: High Quality Video Reasoning Segmentation 1 2 15 Jan 2025 not harvested
VISA: Reasoning Video Object Segmentation via Large Language Models 2 2 16 Jul 2024 ran 2 of 2 samples (0 unverified; 2 pointer-only for licence)
Tracking with Human-Intent Reasoning 1 1 29 Dec 2023 not harvested
MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions 1 1 16 Aug 2023 ran 0 of 5 samples (5 unverified)
LISA: Reasoning Segmentation via Large Language Model 2 1 1 Aug 2023 not harvested
Language as Queries for Referring Video Object Segmentation 1 1 3 Jan 2022 ran 7 of 8 samples (1 unverified; 8 pointer-only for licence)
End-to-End Referring Video Object Segmentation with Multimodal Transformers 2 1 29 Nov 2021 ran 6 of 11 samples (5 unverified)

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

CC BY-NC-SA 4.0

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • ReVOS

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections