Browse State-of-the-Art › Video Panoptic Segmentation
Video Panoptic Segmentation
21 papers with code · 5 benchmarks · 6 datasets archive 2025-07-28
Video Panoptic Segmentation is a computer vision task that extends panoptic segmentation by incorporating temporal dimension. That is, given a video sequence, the goal is to predict the semantic class of each pixel while consistently tracking object instances. Here, the pixels belonging to the same object instance should be assigned the same instance ID throughout the video sequence.
Description from the archive archive 2025-07-28; Papers-with-Code links inside it are rewritten to this site.
Benchmarks archive 2025-07-28
5 leaderboard tables shown for this task, 5 with rows (a “benchmark” on this site is a table with at least one row, as on /sota), ordered by row count. “Best model” is the first row in the archive's own order at snapshot; nothing is re-ranked here and metric direction is not recorded in the archive. PwC's Trend sparklines are not in the archive, so that column is omitted.
| Dataset | Best model (first row in archive order) | Paper | Code | Syntology | Compare |
|---|---|---|---|---|---|
| VIPSeg (12 rows) | CAVIS(VIT-L) | Context-Aware Video Instance Segmentation | code | — | Compare |
| Cityscapes-VPS (8 rows) | VIP-Deeplab | ViP-DeepLab: Learning Visual Perception with Depth-aware Video... | code | Syntology ran 3 of 3 samples · 0 unverified | Compare |
| KITTI-STEP (6 rows) | Video K-Net (Swin-L) | Video K-Net: A Simple, Strong, and Unified Baseline for Video Segmentation | code | — | Compare |
| 4D-OR (2 rows) | MM-OR-VPQ4 | MM-OR: A Large Multimodal Operating Room Dataset for Semantic... | code | — | Compare |
| MM-OR (2 rows) | MM-OR-VPQ4 | MM-OR: A Large Multimodal Operating Room Dataset for Semantic... | code | — | Compare |
Syntology column: samples harvested from the paper's repositories and executed on synthesized fixtures; “ran” is not a correctness claim and does not order the table. A dash means no Syntology record for that paper, not a recorded non-run. Read from the graph 2026-09-24.
Libraries
Not in the archive: the export carries no per-task library table, so there is nothing to show at snapshot 2025-07-28.
Datasets archive 2025-07-28
6 datasets whose archive record lists this task, ordered by the archive's paper count.
Subtasks archive 2025-07-28
No subtask under this task in the archive's task tree.
Parent tasks archive 2025-07-28
Most implemented papers archive 2025-07-28
21 shown of 21 papers with code (42 tagged with this task in all), ordered by repositories listed in the archive, not by stars (the archive holds no stars, so PwC's “Social” and “Latest” sorts cannot be reproduced). Papers without a page here are shown as plain text.
-
30 Nov 2023 2 repositories listedIn this work, we present Axial-VS, a general and simple framework that enhances video segmenters by tracking objects along axial trajectories.
-
9 Dec 2020 2 repositories listed Syntology ran 3 of 3 samples · 0 unverified · 3 pointer-only (licence)We name this joint task as Depth-aware Video Panoptic Segmentation, and propose a new evaluation metric along with two derived datasets for it, which will be made available to the public.
-
4 Mar 2025 1 repository listedOperating rooms (ORs) are complex, high-stakes environments requiring precise understanding of interactions among medical staff, tools, and equipment for enhancing surgical assistance, situational awareness, and patient…
-
3 Jul 2024 1 repository listedIn this paper, we introduce the Context-Aware Video Instance Segmentation (CAVIS), a novel framework designed to enhance instance association by integrating contextual information adjacent to each object.
-
1 Jul 2024 1 repository listedWe present Uni-DVPS, a unified model for Depth-aware Video Panoptic Segmentation (DVPS) that jointly tackles distinct vision tasks, i.
-
16 May 2024 1 repository listedThe findings of the conducted quantitative and qualitative evaluations demonstrate the ability of our framework to spot the most and least influential fragments and visual objects of the video for the summarizer, and to…
-
28 Feb 2024 1 repository listed Syntology ran 12 of 14 samples · 2 unverified · 14 pointer-only (licence)Despite the recent advances in unified image segmentation (IS), developing a unified video segmentation (VS) model remains a challenge.
-
20 Dec 2023 1 repository listedWe present the \textbf{D}ecoupled \textbf{VI}deo \textbf{S}egmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video…
-
7 Sep 2023 1 repository listed Syntology ran 7 of 10 samples · 3 unverified · 10 pointer-only (licence)To 'track anything' without training on video data for every individual task, we develop a decoupled video segmentation approach (DEVA), composed of task-specific image-level segmentation and class/task-agnostic…
-
7 Jun 2023 1 repository listedIn this report, we successfully validated the effectiveness of the decoupling strategy in video panoptic segmentation.
-
6 Jun 2023 1 repository listedThe efficacy of the decoupling strategy relies on two crucial elements: 1) attaining precise long-term alignment outcomes via frame-by-frame association during tracking, and 2) the effective utilization of temporal…
-
22 Mar 2023 1 repository listedOur framework is a near-online approach that takes a short subclip as input and outputs the corresponding spatial-temporal tube masks.
-
6 Jan 2023 1 repository listed Syntology ran 0 of 6 samples · 6 unverifiedA single TarViS model can be trained jointly on a collection of datasets spanning different tasks, and can hot-swap between tasks during inference without any task-specific retraining.
-
1 Jan 2023 1 repository listedWe evaluate the proposed approach across three challenging tasks: video instance segmentation, multi-object tracking and segmentation, and video panoptic segmentation.
-
4 Jul 2022 1 repository listedWe present PVO, a novel panoptic visual odometry framework to achieve more comprehensive modeling of the scene motion, geometry, and panoptic segmentation information.
-
15 Jun 2022 1 repository listedWe therefore present the Waymo Open Dataset: Panoramic Video Panoptic Segmentation Dataset, a large-scale dataset that offers high-quality panoptic segmentation labels for autonomous driving.
-
10 Apr 2022 1 repository listedWe hope this simple, yet effective method can serve as a new, flexible baseline in unified video segmentation design.
-
1 Jan 2022 1 repository listedIn contrast, our large-scale VIdeo Panoptic Segmentation in the Wild (VIPSeg) dataset provides 3, 536 videos and 84, 750 frames with pixel-level panoptic annotations, covering a wide range of real-world scenarios and…
-
5 Dec 2021 1 repository listedThe Depth-aware Video Panoptic Segmentation (DVPS) is a new challenging vision problem that aims to predict panoptic segmentation and depth in a video simultaneously.
-
23 Feb 2021 1 repository listedThe task of assigning semantic classes and track identities to every pixel in a video is called video panoptic segmentation.
-
19 Jun 2020 1 repository listedIn this paper, we propose and explore a new video extension of this task, called video panoptic segmentation.
Syntology lines on 4 of the papers shown; no Syntology record for the others (a paper without an arXiv id cannot be joined to the graph, and absence from the graph layer is not a recorded non-run). “Ran” means the sample executed on a synthesized fixture, not that the paper's result was reproduced. Read from the graph 2026-09-24.
Report a problem or propose a change · a person checks every report against the paper or source before anything changes; decisions are listed on /corrections