{"url":"/method/mivos","slug":"mivos","name":"MiVOS","full_name":"Modular Interactive VOS","full_name_withheld":false,"description_markdown":"**MiVOS** is a video object segmentation model which decouples interaction-to-mask and mask propagation. By decoupling interaction from propagation, MiVOS is versatile and not limited by the type of interactions. It uses three modules: Interaction-to-Mask, Propagation and Difference-Aware Fusion. Trained separately, the interaction module converts user interactions to an object mask, which is then temporally propagated by our propagation module using a novel top-filtering strategy in reading the space-time memory. To effectively take the user's intent into account, a novel difference-aware module is proposed to learn how to properly fuse the masks before and after each interaction, which are aligned with the target frames by employing the space-time memory.","description_state":"present","introduced_year":null,"introduced_by":{"title":"Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware Fusion","paper":"/paper/modular-interactive-video-object-segmentation","first_author":"Ho Kei Cheng","n_authors":3,"url_abs":null,"archive_paper_url":"https://paperswithcode.com/paper/modular-interactive-video-object-segmentation"},"source":{"url":"https://arxiv.org/abs/2103.07941v3","title":"Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware Fusion","url_on_a_paper_host":true},"code_snippet_url":null,"code_snippet_url_on_a_code_host":false,"categories":[{"area":"Computer Vision","area_id":"computer-vision","collection":"Video Object Segmentation Models","url":"/methods/category/video-object-segmentation-models","pwc_aliases":[]}],"n_papers_tagged":1,"archive_num_papers":1,"papers_newest_first":[{"paper":"/paper/modular-interactive-video-object-segmentation","title":"Modular Interactive Video Object Segmentation: Interaction-to-Mask, Propagation and Difference-Aware Fusion","date":"2021-03-14","arxiv_id":"2103.07941","n_code_links":5,"syntology":{"ran":9,"of":18,"unverified":9,"pointer_only":0}}],"papers_shown":1,"tasks":[{"task":"/task/interactive-video-object-segmentation","name":"Interactive Video Object Segmentation","papers":1},{"task":"/task/semantic-segmentation","name":"Semantic Segmentation","papers":1},{"task":"/task/semi-supervised-video-object-segmentation","name":"Semi-Supervised Video Object Segmentation","papers":1},{"task":"/task/video-object-segmentation","name":"Video Object Segmentation","papers":1},{"task":"/task/video-semantic-segmentation","name":"Video Semantic Segmentation","papers":1}],"tasks_shown":5,"n_tasks":5,"usage_by_year":[{"year":"2021","papers":1}],"row_source":"methods_table","archive":{"source":"pwc-archive (Hugging Face), CC BY-SA 4.0","snapshot":"2025-07-28","archive_url":"https://paperswithcode.com/method/mivos"},"syntology_read_at":"2026-09-24T18:15:14+00:00"}