{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/stem-seg-spatio-temporal-embeddings-for","title":"STEm-Seg: Spatio-temporal Embeddings for Instance Segmentation in Videos","arxiv_id":"2003.08429","date":"2020-03-18","proceeding":"ECCV 2020 8","authors":["Ali Athar","Sabarinath Mahadevan","Aljoša Ošep","Laura Leal-Taixé","Bastian Leibe"],"abstract":"Existing methods for instance segmentation in videos typically involve multi-stage pipelines that follow the tracking-by-detection paradigm and model a video clip as a sequence of images. Multiple networks are used to detect objects in individual frames, and then associate these detections over time. Hence, these methods are often non-end-to-end trainable and highly tailored to specific tasks. In this paper, we propose a different approach that is well-suited to a variety of tasks involving instance segmentation in videos. In particular, we model a video clip as a single 3D spatio-temporal volume, and propose a novel approach that segments and tracks instances across space and time in a single stage. Our problem formulation is centered around the idea of spatio-temporal embeddings which are trained to cluster pixels belonging to a specific object instance over an entire video clip. To this end, we introduce (i) novel mixing functions that enhance the feature representation of spatio-temporal embeddings, and (ii) a single-stage, proposal-free network that can reason about temporal context. Our network is trained end-to-end to learn spatio-temporal embeddings as well as parameters required to cluster these embeddings, thus simplifying inference. Our method achieves state-of-the-art results across multiple datasets and tasks. Code and models are available at https://github.com/sabarim/STEm-Seg.","url_abs":"https://arxiv.org/abs/2003.08429v4","url_pdf":"https://arxiv.org/pdf/2003.08429v4.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"stem-seg-spatio-temporal-embeddings-for","repo_url":"https://github.com/sabarim/STEm-Seg","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":1,"framework":"pytorch","reach":{"status":"ok"}}],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"unsupervised-video-object-segmentation","task_name":"Unsupervised Video Object Segmentation"},{"task_slug":"video-instance-segmentation","task_name":"Video Instance Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/unsupervised-video-object-segmentation-on-4","task":"Unsupervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"STEm-Seg","rank_in_archive_order":5,"of":10,"metrics":{"F-measure (Mean)":"67.8","F-measure (Recall)":"75.5","J&F":"64.7","Jaccard (Mean)":"61.5","Jaccard (Recall)":"70.4"},"uses_additional_data":true},{"leaderboard":"/sota/video-instance-segmentation-on-youtube-vis-1","task":"Video Instance Segmentation","dataset":"YouTube-VIS validation","model":"STEm-Seg (ResNet-101)","rank_in_archive_order":35,"of":44,"metrics":{"AP50":"55.8","AP75":"37.9","AR1":"34.4","AR10":"41.6","mask AP":"34.6"},"uses_additional_data":false},{"leaderboard":"/sota/video-instance-segmentation-on-youtube-vis-1","task":"Video Instance Segmentation","dataset":"YouTube-VIS validation","model":"STEm-Seg (ResNet-50)","rank_in_archive_order":40,"of":44,"metrics":{"AP50":"50.7","AP75":"37.9","AR1":"34.4","AR10":"41.6","mask AP":"30.6"},"uses_additional_data":false}],"syntology":{"syntology_url":"https://syntology.ai/paper/2003.08429","atlas_url":"https://app.syntology.ai/?focus=2003.08429","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}