{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/low-latency-video-semantic-segmentation","title":"Low-Latency Video Semantic Segmentation","arxiv_id":"1804.00389","date":"2018-04-02","proceeding":"CVPR 2018 6","authors":["Yule Li","Jianping Shi","Dahua Lin"],"abstract":"Recent years have seen remarkable progress in semantic segmentation. Yet, it\nremains a challenging task to apply segmentation techniques to video-based\napplications. Specifically, the high throughput of video streams, the sheer\ncost of running fully convolutional networks, together with the low-latency\nrequirements in many real-world applications, e.g. autonomous driving, present\na significant challenge to the design of the video segmentation framework. To\ntackle this combined challenge, we develop a framework for video semantic\nsegmentation, which incorporates two novel components: (1) a feature\npropagation module that adaptively fuses features over time via spatially\nvariant convolution, thus reducing the cost of per-frame computation; and (2)\nan adaptive scheduler that dynamically allocate computation based on accuracy\nprediction. Both components work together to ensure low latency while\nmaintaining high segmentation quality. On both Cityscapes and CamVid, the\nproposed framework obtained competitive performance compared to the state of\nthe art, while substantially reducing the latency, from 360 ms to 119 ms.","url_abs":"http://arxiv.org/abs/1804.00389v1","url_pdf":"http://arxiv.org/pdf/1804.00389v1.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"autonomous-driving","task_name":"Autonomous Driving"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-segmentation","task_name":"Video Segmentation"},{"task_slug":"video-semantic-segmentation","task_name":"Video Semantic Segmentation"}],"methods":[],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-semantic-segmentation-on-cityscapes-val","task":"Video Semantic Segmentation","dataset":"Cityscapes val","model":"LVS [12]","rank_in_archive_order":6,"of":9,"metrics":{"mIoU":"76.8"},"uses_additional_data":false}],"syntology":{"atlas_url":"https://app.syntology.ai/?focus=1804.00389","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}