Datasets › YouTube-VIS 2021

YouTube-VIS 2021 (Video Instance Segmentation on YouTube-VIS 2021 validation)

Introduced by Linjie Yang et al. in Video Instance Segmentation12 May 2019 archive 2025-07-28

3,859 high-resolution YouTube videos, 2,985 training videos, 421 validation videos and 453 test videos. An improved 40-category label set by merging eagle and owl into bird, ape into monkey, deleting hands, and adding flying disc, squirrel and whale 8,171 unique video instances 232k high-quality manual annotations

Benchmarks archive 2025-07-28

All 1 leaderboard whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.

First row (archive order)PaperCode
Video Instance Segmentation YouTube-VIS 2021 CAVIS(VIT-L, Offline) mask AP 65.3 Context-Aware Video Instance Segmentation Seung-Hun-Lee/CAVIS 26 Compare

Papers archive 2025-07-28

19 shown of 19 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 52. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.

DateSamples run Syntology
Context-Aware Video Instance Segmentation 1 1 3 Jul 2024 not harvested
DVIS-DAQ: Improving Video Segmentation via Dynamic Anchor Queries 3 1 29 Mar 2024 ran 1 of 5 samples (4 unverified; 3 pointer-only for licence)
UniVS: Unified and Universal Video Segmentation with Prompts as Queries 1 1 28 Feb 2024 ran 12 of 14 samples (2 unverified; 14 pointer-only for licence)
DVIS++: Improved Decoupled Framework for Universal Video Segmentation 1 2 20 Dec 2023 not harvested
NOVIS: A Case for End-to-End Near-Online Video Instance Segmentation 0 2 29 Aug 2023 not harvested
RefineVIS: Video Instance Segmentation with Temporal Attention Refinement 0 1 7 Jun 2023 not harvested
DVIS: Decoupled Video Instance Segmentation Framework 1 1 6 Jun 2023 not harvested
GRAtt-VIS: Gated Residual Attention for Auto Rectifying Video Instance Segmentation 1 2 26 May 2023 not harvested
BoxVIS: Video Instance Segmentation with Box Annotations 1 1 26 Mar 2023 not harvested
MDQE: Mining Discriminative Query Embeddings to Segment Occluded Instances on Challenging Videos 1 1 25 Mar 2023 ran 3 of 4 samples (1 unverified; 4 pointer-only for licence)
Tube-Link: A Flexible Cross Tube Framework for Universal Video Segmentation 1 1 22 Mar 2023 not harvested
TarViS: A Unified Approach for Target-based Video Segmentation 1 3 6 Jan 2023 ran 0 of 6 samples (6 unverified)
A Generalized Framework for Video Instance Segmentation 1 1 16 Nov 2022 not harvested
InstanceFormer: An Online Video Instance Segmentation Framework 1 2 22 Aug 2022 not harvested
MinVIS: A Minimal Video Instance Segmentation Framework without Video-based Training 2 1 3 Aug 2022 not harvested
DeVIS: Making Deformable Transformers Work for Video Instance Segmentation 1 2 22 Jul 2022 ran 2 of 11 samples (9 unverified; 11 pointer-only for licence)
In Defense of Online Models for Video Instance Segmentation 2 1 21 Jul 2022 ran 3 of 5 samples (2 unverified)
VITA: Video Instance Segmentation via Object Token Association 1 1 9 Jun 2022 ran 0 of 7 samples (7 unverified)
Spatial Feature Calibration and Temporal Fusion for Effective One-stage Video Instance Segmentation 1 1 6 Apr 2021 not harvested

Dataset loaders archive 2025-07-28

No loader listed in the archive.

Tasks archive 2025-07-28

License archive 2025-07-28

Creative Commons Attribution 4.0 License

Modalities archive 2025-07-28

Languages archive 2025-07-28

No language tagged.

Variants archive 2025-07-28

  • YouTube-VIS 2021

1 variant name, as the archive lists them.

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections