{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/unified-perception-efficient-video-panoptic","title":"Unified Perception: Efficient Depth-Aware Video Panoptic Segmentation with Minimal Annotation Costs","arxiv_id":"2303.01991","date":"2023-03-03","proceeding":null,"authors":["Kurt Stolle","Gijs Dubbelman"],"abstract":"Depth-aware video panoptic segmentation is a promising approach to camera based scene understanding. However, the current state-of-the-art methods require costly video annotations and use a complex training pipeline compared to their image-based equivalents. In this paper, we present a new approach titled Unified Perception that achieves state-of-the-art performance without requiring video-based training. Our method employs a simple two-stage cascaded tracking algorithm that (re)uses object embeddings computed in an image-based network. Experimental results on the Cityscapes-DVPS dataset demonstrate that our method achieves an overall DVPQ of 57.1, surpassing state-of-the-art methods. Furthermore, we show that our tracking strategies are effective for long-term object association on KITTI-STEP, achieving an STQ of 59.1 which exceeded the performance of state-of-the-art methods that employ the same backbone network. Code is available at: https://tue-mps.github.io/unipercept","url_abs":"https://arxiv.org/abs/2303.01991v2","url_pdf":"https://arxiv.org/pdf/2303.01991v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"depth-aware-video-panoptic-segmentation","task_name":"Depth-aware Video Panoptic Segmentation"},{"task_slug":"panoptic-segmentation","task_name":"Panoptic Segmentation"},{"task_slug":"scene-understanding","task_name":"Scene Understanding"},{"task_slug":"video-panoptic-segmentation","task_name":"Video Panoptic Segmentation"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"},{"method_slug":"fcn","method_name":"FCN"},{"method_slug":"kaiming-initialization","method_name":"Kaiming Initialization"},{"method_slug":"max-pooling","method_name":"Max Pooling"},{"method_slug":"projection-discriminator","method_name":"Projection Discriminator"},{"method_slug":"random-resized-crop","method_name":"Random Resized Crop"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-panoptic-segmentation-on-kitti-step","task":"Video Panoptic Segmentation","dataset":"KITTI-STEP","model":"Unified Perception","rank_in_archive_order":6,"of":6,"metrics":{"AQ":"56.4","SQ":"61.9","STQ":"59.1"},"uses_additional_data":false}],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}