Datasets › ReVOS
ReVOS
We create a benchmark dataset named ReVOS. This dataset comprises 35,074 pairs of instruction-mask sequences derived from 1,042 diverse videos. In contrast to traditional referring video segmentation datasets, such as Ref-YouTube-VOS and MeViS, which primarily contain explicit short phrases, ReVOS includes text instructions that necessitates a sophisticated understanding of both video content and general world knowledge
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Referring Video Object Segmentation | ReVOS | VRS-HQ (Chat-UniVi-13B) J&F 60 | The Devil is in Temporal Token: High Quality Video... | sitonggong/vrs-hq | 9 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 18. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| The Devil is in Temporal Token: High Quality Video Reasoning Segmentation | 1 | 2 | 15 Jan 2025 | not harvested |
| VISA: Reasoning Video Object Segmentation via Large Language Models | 2 | 2 | 16 Jul 2024 | ran 2 of 2 samples (0 unverified; 2 pointer-only for licence) |
| Tracking with Human-Intent Reasoning | 1 | 1 | 29 Dec 2023 | not harvested |
| MeViS: A Large-scale Benchmark for Video Segmentation with Motion Expressions | 1 | 1 | 16 Aug 2023 | ran 0 of 5 samples (5 unverified) |
| LISA: Reasoning Segmentation via Large Language Model | 2 | 1 | 1 Aug 2023 | not harvested |
| Language as Queries for Referring Video Object Segmentation | 1 | 1 | 3 Jan 2022 | ran 7 of 8 samples (1 unverified; 8 pointer-only for licence) |
| End-to-End Referring Video Object Segmentation with Multimodal Transformers | 2 | 1 | 29 Nov 2021 | ran 6 of 11 samples (5 unverified) |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
CC BY-NC-SA 4.0
Modalities archive 2025-07-28
Languages archive 2025-07-28
No language tagged.
Variants archive 2025-07-28
- ReVOS
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections