Datasets › Long-RVOS
Long-RVOS
This work proposes Long-RVOS, a large-scale benchmark for long-term video object segmentation. Long-RVOS is the first minute-level dataset in the RVOS field, designed to tackle various realistic long-video challenges such as frequent occlusion, disappearance-reappearance, and shot changing. Notably, Long-RVOS offers significantly longer video duration than existing datasets. In addition, it contains the largest number of object classes and mask annotations. The large scale of Long-RVOS supports comprehensive training and evaluation of RVOS models. Finally, we gather 24,689 high-quality descriptions for building Long-RVOS.
Benchmarks archive 2025-07-28
All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Referring Video Object Segmentation | Long-RVOS | ReferMo J&F 51.3 | Long-RVOS: A Comprehensive Benchmark for Long-term... | — | 7 | Compare |
Papers archive 2025-07-28
7 shown of 7 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 7. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| Long-RVOS: A Comprehensive Benchmark for Long-term Referring Video Object Segmentation | 0 | 1 | 19 May 2025 | not harvested |
| GLUS: Global-Local Reasoning Unified into A Single Large Language Model for Video Segmentation | 1 | 1 | 10 Apr 2025 | ran 3 of 12 samples (9 unverified; 12 pointer-only for licence) |
| ReferDINO: Referring Video Object Segmentation with Visual Grounding Foundations | 0 | 1 | 24 Jan 2025 | not harvested |
| SAMWISE: Infusing Wisdom in SAM2 for Text-Driven Video Segmentation | 1 | 1 | 26 Nov 2024 | not harvested |
| One Token to Seg Them All: Language Instructed Reasoning Segmentation in Videos | 1 | 1 | 29 Sep 2024 | ran 7 of 16 samples (9 unverified) |
| SOC: Semantic-Assisted Object Cluster for Referring Video Object Segmentation | 1 | 1 | 26 May 2023 | not harvested |
| Referred by Multi-Modality: A Unified Temporal Transformer for Video Object Segmentation | 1 | 1 | 25 May 2023 | not harvested |
Dataset loaders archive 2025-07-28
No loader listed in the archive.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- Long-RVOS
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections