{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/semantic-instance-meets-salient-object-study","title":"Semantic Instance Meets Salient Object: Study on Video Semantic Salient Instance Segmentation","arxiv_id":"1807.01452","date":"2018-07-04","proceeding":null,"authors":["Trung-Nghia Le","Akihiro Sugimoto"],"abstract":"Focusing on only semantic instances that only salient in a scene gains more\nbenefits for robot navigation and self-driving cars than looking at all objects\nin the whole scene. This paper pushes the envelope on salient regions in a\nvideo to decompose them into semantically meaningful components, namely,\nsemantic salient instances. We provide the baseline for the new task of video\nsemantic salient instance segmentation (VSSIS), that is, Semantic Instance -\nSalient Object (SISO) framework. The SISO framework is simple yet efficient,\nleveraging advantages of two different segmentation tasks, i.e. semantic\ninstance segmentation and salient object segmentation to eventually fuse them\nfor the final result. In SISO, we introduce a sequential fusion by looking at\noverlapping pixels between semantic instances and salient regions to have\nnon-overlapping instances one by one. We also introduce a recurrent instance\npropagation to refine the shapes and semantic meanings of instances, and an\nidentity tracking to maintain both the identity and the semantic meaning of\ninstances over the entire video. Experimental results demonstrated the\neffectiveness of our SISO baseline, which can handle occlusions in videos. In\naddition, to tackle the task of VSSIS, we augment the DAVIS-2017 benchmark\ndataset by assigning semantic ground-truth for salient instance labels,\nobtaining SEmantic Salient Instance Video (SESIV) dataset. Our SESIV dataset\nconsists of 84 high-quality video sequences with pixel-wisely per-frame\nground-truth labels.","url_abs":"http://arxiv.org/abs/1807.01452v3","url_pdf":"http://arxiv.org/pdf/1807.01452v3.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"robot-navigation","task_name":"Robot Navigation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"self-driving-cars","task_name":"Self-Driving Cars"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"}],"methods":[],"datasets_introduced":[{"slug":"sesiv","name":"SESIV","full_name":"SEmantic Salient Instance Video"}],"methods_introduced":[],"results":[],"syntology":{"atlas_url":null,"mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}