{"about":{"site":"https://codewithpapers.app","non_affiliation":"Code with Papers and Syntology are not affiliated with, endorsed by, or sponsored by Papers with Code, Meta, or the pwc-archive mirror.","licence":"CC BY-SA 4.0","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","attribution":"https://codewithpapers.app/attribution","modified":"archive material modified by Syntology; see the attribution page"},"url":"/paper/novis-a-case-for-end-to-end-near-online-video","title":"NOVIS: A Case for End-to-End Near-Online Video Instance Segmentation","arxiv_id":"2308.15266","date":"2023-08-29","proceeding":null,"authors":["Tim Meinhardt","Matt Feiszli","Yuchen Fan","Laura Leal-Taixe","Rakesh Ranjan"],"abstract":"Until recently, the Video Instance Segmentation (VIS) community operated under the common belief that offline methods are generally superior to a frame by frame online processing. However, the recent success of online methods questions this belief, in particular, for challenging and long video sequences. We understand this work as a rebuttal of those recent observations and an appeal to the community to focus on dedicated near-online VIS approaches. To support our argument, we present a detailed analysis on different processing paradigms and the new end-to-end trainable NOVIS (Near-Online Video Instance Segmentation) method. Our transformer-based model directly predicts spatio-temporal mask volumes for clips of frames and performs instance tracking between clips via overlap embeddings. NOVIS represents the first near-online VIS approach which avoids any handcrafted tracking heuristics. We outperform all existing VIS methods by large margins and provide new state-of-the-art results on both YouTube-VIS (2019/2021) and the OVIS benchmarks.","url_abs":"https://arxiv.org/abs/2308.15266v2","url_pdf":"https://arxiv.org/pdf/2308.15266v2.pdf","source":{"archive":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","licence_url":"https://creativecommons.org/licenses/by-sa/4.0/legalcode","row_kind":"abstracts"},"code_links":[],"tasks":[{"task_slug":"instance-segmentation","task_name":"Instance Segmentation"},{"task_slug":"segmentation","task_name":"Segmentation"},{"task_slug":"semantic-segmentation","task_name":"Semantic Segmentation"},{"task_slug":"video-instance-segmentation","task_name":"Video Instance Segmentation"}],"methods":[{"method_slug":"focus","method_name":"Focus"}],"datasets_introduced":[],"methods_introduced":[],"results":[{"leaderboard":"/sota/video-instance-segmentation-on-ovis-1","task":"Video Instance Segmentation","dataset":"OVIS validation","model":"NOVIS (Swin-L)","rank_in_archive_order":13,"of":44,"metrics":{"AP50":"68.3","AP75":"43.8","AR1":"19.4","AR10":"46.9","mask AP":"43.5"},"uses_additional_data":true},{"leaderboard":"/sota/video-instance-segmentation-on-ovis-1","task":"Video Instance Segmentation","dataset":"OVIS validation","model":"NOVIS (ResNet-50)","rank_in_archive_order":28,"of":44,"metrics":{"AP50":"56.2","AP75":"32.6","AR1":"15.7","AR10":"37.1","mask AP":"32.7"},"uses_additional_data":true},{"leaderboard":"/sota/video-instance-segmentation-on-youtube-vis-2","task":"Video Instance Segmentation","dataset":"YouTube-VIS 2021","model":"NOVIS (Swin-L)","rank_in_archive_order":10,"of":26,"metrics":{"AP50":"82.0","AP75":"66.5","AR1":"47.9","AR10":"64.4","mask AP":"59.8"},"uses_additional_data":true},{"leaderboard":"/sota/video-instance-segmentation-on-youtube-vis-2","task":"Video Instance Segmentation","dataset":"YouTube-VIS 2021","model":"NOVIS (ResNet-50)","rank_in_archive_order":23,"of":26,"metrics":{"AP50":"69.4","AP75":"50.0","AR1":"41.3","AR10":"54.4","mask AP":"47.2"},"uses_additional_data":true},{"leaderboard":"/sota/video-instance-segmentation-on-youtube-vis-1","task":"Video Instance Segmentation","dataset":"YouTube-VIS validation","model":"NOVIS (ResNet-50)","rank_in_archive_order":14,"of":44,"metrics":{"AP50":"75.7","AP75":"56.9","AR1":"50.3","AR10":"60.6","mask AP":"52.8"},"uses_additional_data":true}],"syntology":{"syntology_url":null,"atlas_url":"https://app.syntology.ai/?focus=2308.15266","mcp":null,"developers":"https://syntology.ai/developers"},"arxiv_metadata":null,"syntology_extracted_results":null}