Papers › Video Panoptic Segmentation

Video Panoptic Segmentation

19 Jun 2020CVPR 2020 6arXiv:2006.11339archive 2025-07-28

Dahun Kim, Sanghyun Woo, Joon-Young Lee, In So Kweon

Panoptic segmentation has become a new standard of visual recognition task by unifying previous semantic segmentation and instance segmentation tasks in concert. In this paper, we propose and explore a new video extension of this task, called video panoptic segmentation. The task requires generating consistent panoptic segmentation as well as an association of instance ids across video frames. To invigorate research on this new task, we present two types of video panoptic datasets. The first is a re-organization of the synthetic VIPER dataset into the video panoptic format to exploit its large-scale pixel annotations. The second is a temporal extension on the Cityscapes val. set, by providing new video panoptic annotations (Cityscapes-VPS). Moreover, we propose a novel video panoptic segmentation network (VPSNet) which jointly predicts object classes, bounding boxes, masks, instance id tracking, and semantic segmentation in video frames. To provide appropriate metrics for this task, we propose a video panoptic quality (VPQ) metric and evaluate our method and several other baselines. Experimental results demonstrate the effectiveness of the presented two datasets. We achieve state-of-the-art results in image PQ on Cityscapes and also in VPQ on Cityscapes-VPS and VIPER datasets. The datasets and code are made publicly available.

PaperPDFConference PDFCode

In Syntology Open this paper in Syntology's Atlas, the map of the papers in Syntology's graph and their citations.

Code

mcahny/vps officialmentioned in papermentioned on GitHubpytorchNOASSERTION report

Repository list and official/mentioned flags are the archive's, frozen 2025-07-28. Reachability, where shown, is from one Syntology probe window (2026-09-16 to 2026-09-18); repositories not probed show nothing. GitHub stars are not tracked.

Code Syntology ran Syntology

Not run by Syntology. Nothing on this page verifies that the listed code works.

Tasks

Instance SegmentationPanoptic SegmentationSegmentationSemantic SegmentationVideo Instance SegmentationVideo Panoptic SegmentationVideo RecognitionVideo SegmentationVideo Semantic Segmentation

Datasets

Introduced by this paper, per the archive.

Cityscapes-VPS

Results from the paper archive 2025-07-28

TaskDatasetModelMetricValueRank at snapshotLeaderboardReport
Video Panoptic Segmentation Cityscapes-VPS VPSNet VPQ 57.0 #7 of 8 Archive leaderboard report
Video Panoptic Segmentation Cityscapes-VPS VPSNet VPQ (stuff) 66.0 #7 of 8 Archive leaderboard report
Video Panoptic Segmentation Cityscapes-VPS VPSNet VPQ (thing) 44.7 #7 of 8 Archive leaderboard report

Ranks are positions in the archive's leaderboards as they stood at the 2025-07-28 snapshot. Results published since then are not among these rows, so a rank here is not a current standing.

Methods

Introduced by this paper: VPSNet

VPSNet

Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections