Datasets › UVO
UVO (Unidentified Video Objects: A Benchmark for Dense, Open-World Segmentation)
UVO is a new benchmark for open-world class-agnostic object segmentation in videos. Besides shifting the problem focus to the open-world setup, UVO is significantly larger, providing approximately 8 times more videos compared with DAVIS, and 7 times more mask (instance) annotations per video compared with YouTube-VOS and YouTube-VIS. UVO is also more challenging as it includes many videos with crowded scenes and complex background motions. Some highlights of the dataset include:
-
High quality instance masks densely annotated at 30 fps on 1024 YouTube videos and 1fps on 10337 videos from Kinetics dataset
-
Open-world: annotating all objects in each video, 13.5 objects per video on average
-
Diverse object categories: 57% of objects are not covered by COCO categories
Benchmarks archive 2025-07-28
All 2 leaderboards whose dataset resolves to this page shown (sort by any header). "First row" is the archive's own first row at snapshot, in the archive's row order; nothing here re-ranks and metric direction is not asserted.
| First row (archive order) | Paper | Code | ||||
|---|---|---|---|---|---|---|
| Open-World Instance Segmentation | UVO | GLEE-Pro ARmask 72.6 | General Object Foundation Model for Images and Videos at Scale | FoundationVision/GLEE | 2 | Compare |
| Unsupervised Instance Segmentation | UVO | CutLER (Cascade+DINO) AP 10.1 | Cut and Learn for Unsupervised Object Detection and... | facebookresearch/cutler +1 | 1 | Compare |
Papers archive 2025-07-28
3 shown of 3 papers with a leaderboard row on this dataset's benchmarks, newest first. The archive's own "papers using this dataset" list was never published, so this is the benchmark-backed subset; the archive's count for this dataset is 27. The Syntology column is from Syntology's graph (read 2026-09-24), stated per sample; it is not part of any archive number.
| Date | Samples run Syntology | |||
|---|---|---|---|---|
| General Object Foundation Model for Images and Videos at Scale | 1 | 1 | 14 Dec 2023 | ran 8 of 13 samples (5 unverified) |
| Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models | 1 | 1 | 8 Mar 2023 | not harvested |
| Cut and Learn for Unsupervised Object Detection and Instance Segmentation | 2 | 1 | 26 Jan 2023 | not harvested |
Dataset loaders archive 2025-07-28
1 loader as listed in the archive; links are outbound and not re-checked here.
Tasks archive 2025-07-28
License archive 2025-07-28
No licence recorded in the archive. Absence here is not a statement about the dataset's terms.
Modalities archive 2025-07-28
Languages archive 2025-07-28
Variants archive 2025-07-28
- UVO
1 variant name, as the archive lists them.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections