{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/adaptive-temporal-encoding-network-for-video","title":"Adaptive Temporal Encoding Network for Video Instance-level Human Parsing","arxiv_id":"1808.00661","date":"2018-08-02","proceeding":null,"authors":["Qixian Zhou","Xiaodan Liang","Ke Gong","Liang Lin"],"abstract":"Beyond the existing single-person and multiple-person human parsing tasks in\nstatic images, this paper makes the first attempt to investigate a more\nrealistic video instance-level human parsing that simultaneously segments out\neach person instance and parses each instance into more fine-grained parts\n(e.g., head, leg, dress). We introduce a novel Adaptive Temporal Encoding\nNetwork (ATEN) that alternatively performs temporal encoding among key frames\nand flow-guided feature propagation from other consecutive frames between two\nkey frames. Specifically, ATEN first incorporates a Parsing-RCNN to produce the\ninstance-level parsing result for each key frame, which integrates both the\nglobal human parsing and instance-level human segmentation into a unified\nmodel. To balance between accuracy and efficiency, the flow-guided feature\npropagation is used to directly parse consecutive frames according to their\nidentified temporal consistency with key frames. On the other hand, ATEN\nleverages the convolution gated recurrent units (convGRU) to exploit temporal\nchanges over a series of key frames, which are further used to facilitate the\nframe-level instance-level parsing. By alternatively performing direct feature\npropagation between consistent frames and temporal encoding network among key\nframes, our ATEN achieves a good balance between frame-level accuracy and time\nefficiency, which is a common crucial problem in video object segmentation\nresearch. To demonstrate the superiority of our ATEN, extensive experiments are\nconducted on the most popular video segmentation benchmark (DAVIS) and a newly\ncollected Video Instance-level Parsing (VIP) dataset, which is the first video\ninstance-level human parsing dataset comprised of 404 sequences and over 20k\nframes with instance-level and pixel-wise annotations.","url_abs":"http://arxiv.org/abs/1808.00661v2","url_pdf":"http://arxiv.org/pdf/1808.00661v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[{"paper_slug":"adaptive-temporal-encoding-network-for-video","repo_url":"https://github.com/HCPLab-SYSU/ATEN","is_official":1,"mentioned_in_paper":1,"mentioned_in_github":0,"framework":"tf","reach":null}],"tasks":[{"task_slug":"human-parsing","task_name":"Human Parsing"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-object-segmentation","task_name":"Video Object Segmentation"},{"task_slug":"video-segmentation","task_name":"Video Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[{"method_slug":"convolution","method_name":"Convolution"}],"datasets_introduced":[],"methods_introduced":[],"results":[],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1808.00661","mcp":{"get_harvested_code_for_paper":{"arxiv_id":"1808.00661"}},"developers":"https://syntology.ai/developers","read_at":"2026-09-24T18:15:14+00:00","read_at_is":"when the build read Syntology's graph, not when any sample ran","claim":"Per-sample execution status on synthesized fixtures; not a correctness claim about the paper. Samples come from repositories linked to the paper, official or community; repo_kind says which.","repos":[{"provenance":"external:paperswithcode_snapshot_2025-07-28","url":"https://github.com/HCPLab-SYSU/ATEN","reach":null}],"summary":{"ran_fixture":1,"unverified":1},"by_repo_kind":{"official":{"samples":1,"ran":0,"repositories":1}},"repo_kind_vocabulary":{"official":"The archive marks this repository official for the paper","named_in_paper":"The archive records that the paper mentions this repository; it is not marked official","listed":"In the archive's code links for this paper, not marked official and not recorded as mentioned in the paper","found_in_text":"Syntology found this repository in the paper's own text; whether it is the authors' implementation is not asserted","community":"Not in the archive's code links for this paper; a community repository Syntology harvested"},"n_pointer_only_for_licence":2,"samples":[{"code_sha256_prefix":"d86fb168246b16a4","entry":"fast_hist","repo":null,"repo_kind":null,"path":null,"file_url":null,"link_basis":"identical_code_first_harvested_elsewhere","language":"python","status":"ran_fixture","verification_level":1,"contract_check":"RAISES","metamorphic_tier":"invariant","behaviour_fingerprint":false,"licence":null,"inline_ok":false,"mcp_get_code":{"code_sha256":"d86fb168246b16a4"}},{"code_sha256_prefix":"384eac2f5e92af52","entry":"compute_hist","repo":"HCPLab-SYSU/ATEN","repo_kind":"official","path":"evaluate/test_parsing.py","file_url":"https://github.com/HCPLab-SYSU/ATEN/blob/HEAD/evaluate/test_parsing.py","link_basis":"first_harvest_node","language":"python","status":"unverified","verification_level":0,"contract_check":null,"metamorphic_tier":null,"behaviour_fingerprint":false,"licence":"NOASSERTION","inline_ok":false,"mcp_get_code":{"code_sha256":"384eac2f5e92af52"}}]},"arxiv_metadata":null,"syntology_extracted_results":null}