{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/spatiotemporal-cnn-for-video-object","title":"Spatiotemporal CNN for Video Object Segmentation","arxiv_id":"1904.02363","date":"2019-04-04","proceeding":"CVPR 2019 6","authors":["Kai Xu","Longyin Wen","Guorong Li","Liefeng Bo","Qingming Huang"],"abstract":"In this paper, we present a unified, end-to-end trainable spatiotemporal CNN\nmodel for VOS, which consists of two branches, i.e., the temporal coherence\nbranch and the spatial segmentation branch. Specifically, the temporal\ncoherence branch pretrained in an adversarial fashion from unlabeled video\ndata, is designed to capture the dynamic appearance and motion cues of video\nsequences to guide object segmentation. The spatial segmentation branch focuses\non segmenting objects accurately based on the learned appearance and motion\ncues. To obtain accurate segmentation results, we design a coarse-to-fine\nprocess to sequentially apply a designed attention module on multi-scale\nfeature maps, and concatenate them to produce the final prediction. In this\nway, the spatial segmentation branch is enforced to gradually concentrate on\nobject regions. These two branches are jointly fine-tuned on video segmentation\nsequences in an end-to-end manner. Several experiments are carried out on three\nchallenging datasets (i.e., DAVIS-2016, DAVIS-2017 and Youtube-Object) to show\nthat our method achieves favorable performance against the state-of-the-arts.\nCode is available at https://github.com/longyin880815/STCNN.","url_abs":"http://arxiv.org/abs/1904.02363v1","url_pdf":"http://arxiv.org/pdf/1904.02363v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"spatiotemporal-cnn-for-video-object","repo_url":"https://github.com/longyin880815/STCNN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"pytorch","reach":null}],"tasks":[{"task_slug":"object","task_name":"Object"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"semi-supervised-video-object-segmentation","task_name":"Semi-Supervised Video Object Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-segmentation","task_name":"Video Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"},{"task_slug":"visual-object-tracking","task_name":"Visual Object Tracking"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/semi-supervised-video-object-segmentation-on-20","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS (no YouTube-VOS training)","model":"STCNN","rank_in_archive_order":25,"of":26,"metrics":{"D16 val (F)":"83.8","D16 val (G)":"83.8","D16 val (J)":"83.8","D17 val (F)":"64.6","D17 val (G)":"61.7","D17 val (J)":"58.7","FPS":"0.26"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2016","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2016","model":"Spatiotemporal CNN","rank_in_archive_order":53,"of":78,"metrics":{"F-measure (Mean)":"83.8","J&F":"83.8","Jaccard (Mean)":"83.8"},"uses_additional_data":false},{"leaderboard":"/sota/visual-object-tracking-on-davis-2017","task":"Semi-Supervised Video Object Segmentation","dataset":"DAVIS 2017 (val)","model":"Spatiotemporal CNN","rank_in_archive_order":70,"of":81,"metrics":{"F-measure (Mean)":"64.6","J&F":"61.65","Jaccard (Mean)":"58.7"},"uses_additional_data":false},{"leaderboard":"/sota/video-object-segmentation-on-youtube","task":"Semi-Supervised Video Object Segmentation","dataset":"YouTube","model":"Spatiotemporal CNN","rank_in_archive_order":1,"of":5,"metrics":{"mIoU":"0.796"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1904.02363","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1904.02363"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/longyin880815/STCNN","reach":null}],"summary":{"ran_violates":1},"by_repo_kind":{"official":{"samples":1,"ran":1,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":1,"samples":[{"code_sha256_prefix":"e47297b1db7138e9","entry":"inverse_transform","repo":"longyin880815/STCNN","repo_kind":"official","path":"train_JointModel.py","file_url":"https://github.com/longyin880815/STCNN/blob/HEAD/train_JointModel.py","link_basis":"first_harvest_node","language":"python","status":"ran_violates","verification_level":1,"contract_check":"VIOLATES","metamorphic_tier":"invariant","behaviour_fingerprint":true,"licence":"NONE","inline_ok":false,"mcp_get_code":{"code_sha256":"e47297b1db7138e9"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}